clastogen 0.2.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,44 @@
1
+ # Bytecode and runtime caches
2
+ __pycache__/
3
+ *.py[cod]
4
+ *$py.class
5
+
6
+ # Packaging and builds
7
+ build/
8
+ dist/
9
+ wheels/
10
+ *.egg-info/
11
+ .eggs/
12
+
13
+ # Virtual environments
14
+ .venv/
15
+ env/
16
+ venv/
17
+
18
+ # Testing, typing, and linting
19
+ .pytest_cache/
20
+ .mypy_cache/
21
+ .ruff_cache/
22
+ .hypothesis/
23
+
24
+ # Coverage
25
+ .coverage
26
+ .coverage.*
27
+ coverage.xml
28
+ htmlcov/
29
+
30
+ # Local environment and secrets
31
+ .env
32
+ .env.*
33
+ !.env.example
34
+
35
+ # IDE and OS artifacts
36
+ .DS_Store
37
+ Thumbs.db
38
+ .idea/
39
+ .vscode/
40
+ *.swp
41
+ *.swo
42
+
43
+ # Clastogen report output (--clastogen-json / --clastogen-html)
44
+ reports/
@@ -0,0 +1,52 @@
1
+ # CHANGELOG
2
+
3
+ <!-- version list -->
4
+
5
+ ## v0.2.0 (2026-10-05)
6
+
7
+ ### Bug Fixes
8
+
9
+ - **mutator**: Spread mutants across rules and operators
10
+ ([`91433b3`](https://github.com/burakkaygusuz/clastogen/commit/91433b35742f23984375ee003978e5f40eab325c))
11
+
12
+ - **plugin**: Preserve pytest exit status and skip xfailed trials
13
+ ([`1dbea5e`](https://github.com/burakkaygusuz/clastogen/commit/1dbea5e29026bdb0bd618418a5983623f215d277))
14
+
15
+ - **stats**: Reject min_rate of 1.0 in evaluate_pass_rate
16
+ ([`2890f5d`](https://github.com/burakkaygusuz/clastogen/commit/2890f5db64f8253bdd2419faaafe4cf70b3ae425))
17
+
18
+ ### Build System
19
+
20
+ - **sdist**: Limit source distribution to package files
21
+ ([`2f15372`](https://github.com/burakkaygusuz/clastogen/commit/2f1537221aa7fb55fe645b6d596fd7aa37f5771d))
22
+
23
+ ### Continuous Integration
24
+
25
+ - **actions**: Pin actions to commit SHAs and add dependabot
26
+ ([`5908f64`](https://github.com/burakkaygusuz/clastogen/commit/5908f64a85e2b9c7a023f6ead09f1195029e9591))
27
+
28
+ - **release**: Push releases with a GitHub App token
29
+ ([`e0bb90e`](https://github.com/burakkaygusuz/clastogen/commit/e0bb90ecb63c23d39b7cba10aa06de63339eda97))
30
+
31
+ ### Documentation
32
+
33
+ - **readme**: Correct SPRT call counts and document baseline noise
34
+ ([`77086fa`](https://github.com/burakkaygusuz/clastogen/commit/77086fae635c3f9ff156f2ec97ab13414fee1276))
35
+
36
+ ### Refactoring
37
+
38
+ - **core**: Collapse single-module packages into modules
39
+ ([`d5c0f49`](https://github.com/burakkaygusuz/clastogen/commit/d5c0f492daf384d20bf81e3dafd0b5c5c2ceed84))
40
+
41
+ - **plugin**: Read suppressions from TOML with stdlib tomllib
42
+ ([`3402808`](https://github.com/burakkaygusuz/clastogen/commit/3402808b395f071ca71ad5aa6573e65801491927))
43
+
44
+ ### Breaking Changes
45
+
46
+ - **plugin**: Suppressions move from .clastogen/suppressions.yaml to .clastogen/suppressions.toml
47
+ using [[suppressions]] tables, and the clastogen[config] extra is removed.
48
+
49
+
50
+ ## v0.1.0 (2026-10-05)
51
+
52
+ - Initial Release
@@ -0,0 +1,21 @@
1
+ # MIT License
2
+
3
+ Copyright (c) 2026 Burak Kaygusuz
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,308 @@
1
+ Metadata-Version: 2.5
2
+ Name: clastogen
3
+ Version: 0.2.0
4
+ Summary: Mutation testing and statistical assertion framework for LLMs and AI Agents
5
+ Project-URL: Homepage, https://github.com/burakkaygusuz/clastogen
6
+ Project-URL: Repository, https://github.com/burakkaygusuz/clastogen
7
+ Project-URL: Issues, https://github.com/burakkaygusuz/clastogen/issues
8
+ Author: Burak Kaygusuz
9
+ License-Expression: MIT
10
+ License-File: LICENSE
11
+ Keywords: ai-agents,evaluations,llm,mutation-testing,pytest,sprt
12
+ Classifier: Development Status :: 3 - Alpha
13
+ Classifier: Framework :: Pytest
14
+ Classifier: Intended Audience :: Developers
15
+ Classifier: License :: OSI Approved :: MIT License
16
+ Classifier: Programming Language :: Python :: 3
17
+ Classifier: Programming Language :: Python :: 3.11
18
+ Classifier: Programming Language :: Python :: 3.12
19
+ Classifier: Programming Language :: Python :: 3.13
20
+ Classifier: Programming Language :: Python :: 3.14
21
+ Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
22
+ Classifier: Topic :: Software Development :: Testing
23
+ Requires-Python: >=3.11
24
+ Requires-Dist: pytest>=8.4.0
25
+ Description-Content-Type: text/markdown
26
+
27
+ # clastogen 🧬
28
+
29
+ Mutation testing and statistical assertion framework for LLMs and AI Agents.
30
+
31
+ Clastogen injects controlled faults (deleting constraints, inverting rules, changing numeric limits) into your system prompts to verify whether your test suite actually catches prompt breakages, using Sequential Probability Ratio Tests (SPRT) to stop early and save API budget.
32
+
33
+ ---
34
+
35
+ ## Installation
36
+
37
+ ```bash
38
+ # Install directly from GitHub
39
+ pip install git+https://github.com/burakkaygusuz/clastogen.git
40
+
41
+ # Or for local development with uv
42
+ git clone https://github.com/burakkaygusuz/clastogen.git
43
+ cd clastogen
44
+ uv sync --dev
45
+ ```
46
+
47
+ Note: PyPI publication is pending. Use git installation or editable local checkout.
48
+
49
+ ---
50
+
51
+ ## Usage
52
+
53
+ ### 1. Pytest Plugin (Recommended)
54
+
55
+ Mark your LLM evaluation test with `@pytest.mark.clastogen` and point it to the target prompt variable:
56
+
57
+ ```python
58
+ import pytest
59
+
60
+
61
+ @pytest.mark.clastogen(target="my_app.agent:SYSTEM_PROMPT")
62
+ def test_agent_behavior():
63
+ response = call_agent("transfer $500")
64
+ assert "Verification code" in response
65
+ ```
66
+
67
+ Run tests with mutation testing enabled:
68
+
69
+ ```bash
70
+ # Run all tests with mutation testing
71
+ pytest --clastogen
72
+
73
+ # Target a specific test file
74
+ pytest --clastogen tests/test_agent.py
75
+
76
+ # Verbose output with step-by-step logs
77
+ pytest --clastogen -s --log-cli-level=INFO
78
+ ```
79
+
80
+ #### Marker Options
81
+
82
+ ```python
83
+ @pytest.mark.clastogen(
84
+ target="my_app.agent:SYSTEM_PROMPT", # Required: 'module:VAR' or 'module:Class.ATTR'
85
+ max_mutants=5, # Max mutants generated (default: 5)
86
+ delta=0.30, # Minimum detectable pass-rate drop (default: 0.30)
87
+ p0=0.90, # Known baseline pass rate (default: measured over 10 baseline runs; setting it skips them)
88
+ max_steps=20, # Maximum evaluation runs per mutant (default: 20)
89
+ )
90
+ ```
91
+
92
+ #### Prompt must be read at call time
93
+
94
+ Clastogen mutates the target by swapping the module (or class) attribute for the duration of each trial, then restoring it. It also swaps module-level aliases that hold the very same string object under the same name (e.g. `from my_app.agent import SYSTEM_PROMPT` at the top of another module). Your code therefore has to read the prompt when it calls the model. Copies made earlier are never updated, so mutants against them always appear `SURVIVED`.
95
+
96
+ Function-scoped fixtures are torn down and rebuilt inside the mutated prompt on every trial, so a fixture such as `bot = {"system": agent.PROMPT}` sees the mutant and is fine. Module- and session-scoped fixtures are built once and keep the original prompt, as do import-time captures; mutants against them also appear `SURVIVED`.
97
+
98
+ Wrong: the prompt is captured once at import time, so the mutated attribute is never read.
99
+
100
+ ```python
101
+ # my_app/agent.py
102
+ SYSTEM_PROMPT = "You must never approve refunds over $50."
103
+ CONFIG = {"system": SYSTEM_PROMPT} # frozen copy; also true for f-strings, default args, clients built at import
104
+
105
+
106
+ def ask(question: str) -> str:
107
+ return client.chat(system=CONFIG["system"], user=question)
108
+ ```
109
+
110
+ Right: the module attribute is looked up on every call (a `from my_app.agent import SYSTEM_PROMPT` inside the function body works too).
111
+
112
+ ```python
113
+ # my_app/agent.py
114
+ SYSTEM_PROMPT = "You must never approve refunds over $50."
115
+
116
+
117
+ def ask(question: str) -> str:
118
+ return client.chat(system=SYSTEM_PROMPT, user=question)
119
+ ```
120
+
121
+ #### Parallel runs (pytest-xdist)
122
+
123
+ `pytest --clastogen -n 4` works: each worker returns its records with the test report and the controller merges them into one summary. Killed-mutant short-circuiting and baseline measurement are per-process, so every worker measures its own baselines and may re-evaluate a mutant already killed on another worker. Parallel runs trade some redundant work for wall-clock time.
124
+
125
+ ---
126
+
127
+ ### 2. Statistical Assertions (Python API)
128
+
129
+ Assert non-deterministic LLM behavior with statistical guarantees instead of a single flaky `assert`. `assert_pass_rate` samples a zero-argument callable returning `bool` with Wald's SPRT, stopping as soon as the evidence is decisive (at most `max_samples` calls), and raises `AssertionError` when the pass rate is judged below `min_rate` (see [Pass-rate guarantees](#pass-rate-guarantees)). `assert_no_regression` runs a baseline and a candidate callable (paired McNemar test by default, Fisher exact with `paired=False`) and raises only on a statistically significant drop. Both return a result object, and `evaluate_pass_rate` / `evaluate_regression` return it without asserting. `compute_wilson_interval` is exported too.
130
+
131
+ The example below is [`examples/test_stats_usage.py`](examples/test_stats_usage.py); it uses seeded fake evaluators, so it runs with `pytest examples/test_stats_usage.py` and no API key:
132
+
133
+ <!-- example: examples/test_stats_usage.py -->
134
+ ```python
135
+ """Statistical assertions on seeded fake evaluators: zero API keys, deterministic."""
136
+
137
+ import random
138
+ from collections.abc import Callable
139
+
140
+ import pytest
141
+
142
+ from clastogen import assert_no_regression, assert_pass_rate
143
+
144
+
145
+ def fake_llm(pass_rate: float, seed: int) -> Callable[[], bool]:
146
+ rng = random.Random(seed)
147
+ return lambda: rng.random() < pass_rate
148
+
149
+
150
+ def test_order_intent_pass_rate() -> None:
151
+ result = assert_pass_rate(fake_llm(0.97, seed=1), min_rate=0.90, description="order intent")
152
+ assert result.observed_rate >= 0.90
153
+
154
+
155
+ def test_prompt_upgrade_does_not_regress() -> None:
156
+ assert_no_regression(fake_llm(0.90, seed=2), fake_llm(0.90, seed=3), description="prompt v1 -> v2")
157
+
158
+
159
+ def test_broken_prompt_is_detected_as_regression() -> None:
160
+ with pytest.raises(AssertionError, match="regression"):
161
+ assert_no_regression(fake_llm(0.95, seed=4), fake_llm(0.40, seed=5), description="prompt v1 -> v3")
162
+ ```
163
+
164
+ ---
165
+
166
+ ## Demo
167
+
168
+ Zero API keys required. [`examples/mock_agent.py`](examples/mock_agent.py) is a deterministic mock banking agent that obeys exactly the rules its system prompt currently states, and [`examples/test_banking_eval.py`](examples/test_banking_eval.py) holds two evals: a strong identity-verification test and a deliberately weak refund test that only checks the response is non-empty.
169
+
170
+ ```bash
171
+ uv run pytest --clastogen examples/test_banking_eval.py
172
+ ```
173
+
174
+ ```text
175
+ ======================= Clastogen Mutation Testing Summary =======================
176
+ Total Unique Mutants : 5
177
+ Killed (Caught by Suite): 2
178
+ Survived (Blind Spots) : 3
179
+ Inconclusive (Truncated): 0
180
+ Execution Errors (Excl.): 0
181
+ Skipped (Excl.) : 0
182
+ Suppressed (Excl.) : 0
183
+ Mutation Score : 40.0%
184
+ --------------------------------------------------------------------------------
185
+ [2caac4fe522b] ✗ SURVIVED (6 runs, LLR=-2.38) -> Deleted load-bearing constraint: 'You must never approve refund requests exceeding $...'
186
+ [45e77dddc073] ✓ KILLED (2 runs, LLR=+3.05) (killed by test_identity_verification_is_enforced) -> Inverted constraint: 'You must always verify customer identity' -> 'You must NEVER verify customer identity '
187
+ [55063f25e7dd] ✓ KILLED (2 runs, LLR=+3.05) (killed by test_identity_verification_is_enforced) -> Deleted load-bearing constraint: 'You must always verify customer identity before pr...'
188
+ [7d1281020f96] ✗ SURVIVED (6 runs, LLR=-2.38) -> Inverted constraint: 'You must never approve refund requests e' -> 'You must ALWAYS approve refund requests '
189
+ [add892bdcea2] ✗ SURVIVED (6 runs, LLR=-2.38) -> Changed threshold: '50' -> '500' in 'You must never approve refund requests e'
190
+ --------------------- Baselines (Laplace p0 used for SPRT) ---------------------
191
+ examples/test_banking_eval.py::test_identity_verification_is_enforced: 10/10 baseline runs passed, p0=0.917 -> examples.mock_agent:BANKING_SYSTEM_PROMPT
192
+ examples/test_banking_eval.py::test_refund_handling_superficial_eval: 10/10 baseline runs passed, p0=0.917 -> examples.mock_agent:BANKING_SYSTEM_PROMPT
193
+ ================================================================================
194
+ ```
195
+
196
+ Both tests pass in a normal run, yet only 2 of the 5 mutants are caught (**Mutation Score 40.0%**). The strong identity test kills both the deleted and the inverted identity rule (2 runs each), while the weak refund test keeps passing when the refund rule is deleted, inverted to "ALWAYS approve" or its limit raised from $50 to $500, so those three mutants SURVIVE: an eval blind spot you would otherwise ship.
197
+
198
+ ### Reading the report
199
+
200
+ - **Statuses:** `KILLED` (a test failed under the mutant), `SURVIVED`, `INCONCLUSIVE` (SPRT hit `max_steps` without a verdict), `ERROR` (the evaluation itself broke; the reason is in the JSON `error` field), `SKIPPED`, `SUPPRESSED`.
201
+ - **Mutation Score** = `KILLED / (KILLED + SURVIVED + INCONCLUSIVE)`. `ERROR`, `SKIPPED` and `SUPPRESSED` mutants are not measured and stay out of the score. When several tests evaluate the same mutant, the strongest outcome wins (`KILLED` > `SURVIVED` > `INCONCLUSIVE` > `ERROR` > `SKIPPED` > `SUPPRESSED`).
202
+ - **Baselines:** each marked test is first re-run 9 more times unmutated (10 runs in total). p0 is the Laplace estimate `(s + 1) / (n + 2)`, and a test whose raw pass rate is below 80% is rejected as flaky and not mutated. The terminal "Baselines" section lists them.
203
+ - **Flags:** `--clastogen-fail-under MIN_SCORE` fails the run below that score; `--clastogen-json PATH` writes `mutation_score`, `total_mutants`, `counts` (all six statuses), `results` (one merged record per mutant), `executions` (every test's record per mutant, the data behind the HTML kill matrix) and `baselines`; `sample_count` and `llr` are `null` for mutants the SPRT never ran (`ERROR`, `SKIPPED`, `SUPPRESSED`). `--clastogen-html PATH` renders the same data as a self-contained HTML report.
204
+
205
+ ## Suppressing Mutants
206
+
207
+ When a mutant survives because the foundation model inherently obeys the rule from pre-training (an equivalent mutant), suppress it so it does not skew your score. Suppressed mutants are reported as `SUPPRESSED` and excluded from the Mutation Score. Mutant IDs are the 12-character hashes printed in the summary, for example `7d1281020f96` (the inverted refund rule in the [Demo](#demo)).
208
+
209
+ List them in `.clastogen/suppressions.toml` under your pytest `rootdir`. Every entry needs a `reason`:
210
+
211
+ ```toml
212
+ [[suppressions]]
213
+ mutant_id = "7d1281020f96"
214
+ reason = "Model refuses large refunds regardless of the prompt"
215
+ ```
216
+
217
+ A malformed file aborts a `--clastogen` run with a usage error; runs without `--clastogen` never read it.
218
+
219
+ ---
220
+
221
+ ## Empirical SPRT Performance & Limits
222
+
223
+ SPRT dynamically sizes samples using Wald's sequential boundaries. Severe defects halt in as few as 2 calls with a measured 10/10 baseline (3 with an explicit $p_0 = 0.90$), while ambiguous boundaries cap at $N_{\max}$.
224
+
225
+ Results from Monte Carlo simulation (`PYTHONPATH=src uv run python scripts/simulate_sprt.py`, 10,000 trials per row, seed 42):
226
+
227
+ ### Scenario 1: Standard Evaluation ($p_0 = 0.90, p_1 = 0.60, N_{\max} = 20$)
228
+
229
+ | True P | KILLED % | SURVIVED % | INCONCLUSIVE % | ASN (Mean) | Call Range | Note |
230
+ | :--- | :--- | :--- | :--- | :--- | :--- | :--- |
231
+ | **0.95** | 0.28% | 98.84% | 0.88% | **7.54** | [4, 20] | Clean pass |
232
+ | **0.90 ($H_0$)** | **3.10%** | 91.66% | 5.24% | **9.38** | [3, 20] | False Kill $\le \alpha = 5\%$ |
233
+ | **0.85** | 9.81% | 76.70% | 13.49% | **10.94** | [3, 20] | Slight degradation |
234
+ | **0.75** | 39.92% | 39.42% | **20.66%** | **11.90** | [3, 20] | Indifference zone |
235
+ | **0.60 ($H_1$)** | 84.74% | **8.50%** | 6.76% | **9.03** | [3, 20] | False Survive $\le \beta = 10\%$ |
236
+ | **0.40** | 99.37% | 0.51% | 0.12% | **5.36** | [3, 20] | Severe defect |
237
+ | **0.10** | **100.00%** | 0.00% | 0.00% | **3.33** | [3, 9] | Severe defect |
238
+ | **0.00** | **100.00%** | 0.00% | 0.00% | **3.00** | [3, 3] | **Halted in 3 calls** |
239
+
240
+ ### Scenario 2: Flaky Baseline ($p_0 = 0.75, p_1 = 0.50, N_{\max} = 25$)
241
+
242
+ | True P | KILLED % | SURVIVED % | INCONCLUSIVE % | ASN (Mean) | Call Range | Note |
243
+ | :--- | :--- | :--- | :--- | :--- | :--- | :--- |
244
+ | **0.85** | 0.16% | 98.28% | 1.56% | **9.95** | [6, 25] | High pass |
245
+ | **0.75 ($H_0$)** | **2.92%** | 79.95% | 17.13% | **14.63** | [5, 25] | Flaky baseline false kill $\le 5\%$ |
246
+ | **0.65** | 19.16% | 43.53% | **37.31%** | **18.08** | [5, 25] | **High ambiguity / wandering** |
247
+ | **0.50 ($H_1$)** | 71.34% | **7.23%** | 21.43% | **15.91** | [5, 25] | False Survive $\le \beta = 10\%$ |
248
+ | **0.35** | 97.88% | 0.42% | 1.70% | **10.19** | [5, 25] | Severe defect |
249
+ | **0.10** | **100.00%** | 0.00% | 0.00% | **5.69** | [5, 16] | Severe defect |
250
+ | **0.00** | **100.00%** | 0.00% | 0.00% | **5.00** | [5, 5] | Halted at min bound |
251
+
252
+ ### Scenario 3: Full Plugin Flow (10-run Laplace baseline, $\delta = 0.30$, $N_{\max} = 20$)
253
+
254
+ The plugin measures the unmutated test 10 times, estimates $p_0 = (s + 1) / (n + 2)$ (Laplace rule of succession), rejects baselines whose raw pass rate is below 80%, then runs the SPRT on each mutant with $p_1 = p_0 - 0.30$. Rates below are over accepted baselines.
255
+
256
+ | Baseline P | Mutant P | Baseline rejected % | KILLED % | SURVIVED % | INCONCLUSIVE % | ASN (Mean) | Note |
257
+ | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- |
258
+ | **0.95** | 0.95 | 0.92% | **0.38%** | 98.13% | 1.48% | **7.40** | Unchanged mutant |
259
+ | **0.90** | 0.90 | 5.28% | **2.19%** | 92.61% | 5.20% | **8.57** | Unchanged mutant |
260
+ | **0.85** | 0.85 | 14.00% | **5.02%** | 86.05% | 8.93% | **9.58** | Unchanged mutant |
261
+ | **0.95** | 0.60 | 1.00% | 76.61% | 11.05% | 12.34% | **9.64** | Moderate defect |
262
+ | **0.95** | 0.30 | 0.83% | **99.73%** | 0.13% | 0.14% | **4.52** | Severe defect |
263
+ | **0.95** | 0.00 | 0.83% | **100.00%** | 0.00% | 0.00% | **2.43** | Total failure |
264
+
265
+ > **Indifference Zone Tradeoff:** In ambiguous zones where the mutant pass rate is close to $(p_0 + p_1) / 2$, sequential tests cannot make a definitive call without infinite samples. Clastogen caps trials at $N_{\max}$ and honestly classifies borderline outcomes as `INCONCLUSIVE` instead of guessing.
266
+
267
+ ### Pass-rate guarantees
268
+
269
+ `assert_pass_rate` tests $H_0$: rate $=$ `min_rate` against $H_1$: rate $=$ `min_rate - tolerance`, with both error rates set to `1 - confidence`. A true rate at or above `min_rate` passes with probability of about `confidence`, a true rate at or below `min_rate - tolerance` fails with about the same probability, and rates in between (the indifference zone) can go either way. If `max_samples` runs out before a decision, the sign of the log-likelihood ratio decides; `result.decided` tells you whether the SPRT stopped on its own. Narrow `tolerance` for a sharper verdict at the cost of more samples.
270
+
271
+ Scenario 4 of the simulation script, defaults (`min_rate=0.90, tolerance=0.20, confidence=0.95, max_samples=50`):
272
+
273
+ | True P | PASSED % | ASN (Mean) | Note |
274
+ | :--- | :--- | :--- | :--- |
275
+ | **0.95** | 99.74% | **16.59** | Clearly good |
276
+ | **0.90 (`min_rate`)** | **95.53%** | **23.50** | Pass $\ge$ confidence |
277
+ | **0.85** | 76.89% | **29.42** | Indifference zone |
278
+ | **0.80** | 44.22% | **29.46** | Indifference zone |
279
+ | **0.70 (`min_rate - tolerance`)** | **5.94%** | **19.30** | Fail $\approx$ confidence |
280
+ | **0.60** | 0.44% | **11.63** | Clearly bad |
281
+
282
+ ---
283
+
284
+ ## Known Limitations (v0.1)
285
+
286
+ 1. **Mutation Operator Scope:** Rules are found lexically: a sentence is a candidate only if it contains a keyword such as `must`, `never`, `always`, `avoid`, `only`, `may not` or `require`. Rules phrased without one (for example conditionals like "Escalate to a human if …") are not mutated. Three operators run on each rule: `delete_constraint`, `invert_negation` and `change_threshold` (first number ×10). Semantic LLM-guided mutations, RAG context poisoning, and tool schema mutators are planned for v0.2/v0.3.
287
+ 2. **Equivalent Mutants:** A prompt mutation can occasionally result in identical agent behavior (e.g. if the underlying foundation model inherently obeys a safety constraint from pre-training). Clastogen handles this pragmatically via triage suppression (`.clastogen/suppressions.toml`) rather than automated semantic equivalence proofs.
288
+ 3. **Effect Size:** The SPRT looks for an absolute drop of `delta` (default 0.30) from a 10-run Laplace baseline, which caps p0 at 11/12 ≈ 0.917. Small regressions such as 0.95 → 0.88 need roughly 60-70 runs per mutant to decide, so with `max_steps=20` they end `INCONCLUSIVE` or `SURVIVED`. Lower `delta` and raise `max_steps` per test (`@pytest.mark.clastogen(delta=0.10, max_steps=100)`) when such drops matter, at the matching API cost.
289
+ 4. **Baseline Noise:** The first baseline run is the already-passing test and p0 is a point estimate, so for baselines near the 80% acceptance floor the per-mutant false-kill rate exceeds alpha. For an unchanged mutant (`delta=0.30`, `max_steps=20`, among accepted baselines) a true baseline of 0.90 gives 2.2%, 0.85 gives 5.0%, 0.80 gives 8.6% and 0.75 gives 13.2%. Raising the 80% floor does not fix it (with a 90% floor, 0.85 still gives 7.0%): stabilize the test or set p0 explicitly.
290
+ 5. **Differential Execution:** Clastogen deduplicates already-killed mutants across tests, but does not yet construct a pre-execution static dependency graph.
291
+
292
+ ---
293
+
294
+ ## Commands Summary
295
+
296
+ | Command | Purpose |
297
+ | --- | --- |
298
+ | `pytest` | Run normal test suite (clastogen dormant) |
299
+ | `pytest --clastogen` | Run test suite with mutation testing & summary score |
300
+ | `PYTHONPATH=src uv run python scripts/simulate_sprt.py` | Run Monte Carlo SPRT power simulation |
301
+ | `uv run ruff check .` | Run static code analysis & linter |
302
+ | `uv run mypy src tests scripts` | Run strict static type checking |
303
+
304
+ ---
305
+
306
+ ## License
307
+
308
+ MIT License. See [LICENSE](LICENSE) for details.