contamcheck 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Sofie Nguyen
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,152 @@
1
+ Metadata-Version: 2.4
2
+ Name: contamcheck
3
+ Version: 0.1.0
4
+ Summary: Check whether a language model has memorised a benchmark.
5
+ Author: Sofie Nguyen
6
+ License-Expression: MIT
7
+ Project-URL: Homepage, https://github.com/Sofienguy1/contamcheck
8
+ Project-URL: Issues, https://github.com/Sofienguy1/contamcheck/issues
9
+ Keywords: llm,benchmark,contamination,evaluation,memorization,machine-learning
10
+ Classifier: Development Status :: 3 - Alpha
11
+ Classifier: Environment :: Console
12
+ Classifier: Intended Audience :: Science/Research
13
+ Classifier: Intended Audience :: Developers
14
+ Classifier: Programming Language :: Python :: 3
15
+ Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
16
+ Requires-Python: >=3.10
17
+ Description-Content-Type: text/markdown
18
+ License-File: LICENSE
19
+ Requires-Dist: numpy>=1.24
20
+ Requires-Dist: torch>=2.1
21
+ Requires-Dist: transformers>=4.56
22
+ Requires-Dist: datasets>=2.18
23
+ Requires-Dist: tqdm
24
+ Provides-Extra: dev
25
+ Requires-Dist: pytest; extra == "dev"
26
+ Dynamic: license-file
27
+
28
+ # contamcheck
29
+
30
+ **Has your language model memorised the benchmark?**
31
+
32
+ When a model scores 90% on GSM8K, the score only means something if the model hasn't seen the test questions before. Benchmarks are public on GitHub, in papers and in blog posts, so they leak into the web-scale data models are trained on. A model that memorised the answers looks smart without being smart.
33
+
34
+ `contamcheck` tests a model for signs that it has seen a benchmark's questions during training.
35
+
36
+ ```
37
+ $ contamcheck experiments/models/neo-gsm8k-10ep gsm8k
38
+
39
+ contamcheck model: experiments/models/neo-gsm8k-10ep benchmark: gsm8k reference: openai-community/gpt2
40
+ 200 questions, control: number twins (23 skipped)
41
+
42
+ ✗ completion: 45% of questions flagged (70/154) — 5% expected by chance (p < 0.001)
43
+ Model writes the exact numbers from the second half of the question
44
+ ✗ min-k: 53% of questions flagged (106/200) — 5% expected by chance (p < 0.001)
45
+ Even the 20% most surprising words aren't surprising to the model (vs reference model)
46
+
47
+ Likely contaminated — the model has probably seen these questions in training
48
+ ```
49
+
50
+ *(A model deliberately trained on half of the GSM8K questions being tested. See [Does it work?](#does-it-work))*
51
+
52
+ ## Install
53
+
54
+ ```bash
55
+ pip install contamcheck
56
+ ```
57
+
58
+ Runs on Apple Silicon (MPS), NVIDIA GPUs (CUDA) or CPU. Any causal language model on the Hugging Face Hub, or a local folder, works.
59
+
60
+ ## Usage
61
+
62
+ ```bash
63
+ contamcheck Qwen/Qwen2.5-0.5B gsm8k # built-in benchmark
64
+ contamcheck Qwen/Qwen2.5-0.5B gsm8k -n 500 # more questions, more power
65
+ contamcheck my-model/ questions.jsonl --field question # your own model and benchmark
66
+ contamcheck my-model/ hf:openai/gsm8k:main:test:question # any Hugging Face dataset
67
+ contamcheck my-model/ humaneval --control fresh.jsonl # control questions written after the model's cutoff
68
+ contamcheck my-model/ gsm8k --json # machine-readable output
69
+ contamcheck big-model/ gsm8k --reference EleutherAI/pythia-6.9b # choose the reference yourself
70
+ ```
71
+
72
+ The exit code is 1 when contamination is likely, so it can gate a CI pipeline.
73
+
74
+ ## How it works
75
+
76
+ ### The twin trick
77
+
78
+ Each benchmark question gets a **twin**: the same text with its numbers swapped for other numbers people used elsewhere in the benchmark.
79
+
80
+ > **Original:** Janet's ducks lay **16** eggs per day. She eats **3** for breakfast...
81
+ > **Twin:** Janet's ducks lay **24** eggs per day. She eats **5** for breakfast...
82
+
83
+ A model that never saw the benchmark has no reason to prefer either version. A model that memorised it prefers the original, because those are the numbers it saw. Each test measures that preference for every question.
84
+
85
+ For a clean model, the preferences land on both sides of zero, and the negative side shows what chance looks like on the positive side. A question is flagged when its preference is stronger than 95% of that mirror image. About 5% of questions get flagged by chance. Far more than that, confirmed by a sign-flip permutation test, means the model has seen the benchmark.
86
+
87
+ ### The tests
88
+
89
+ - **completion**: give the model the first half of the question and check whether it writes the **numbers** of the second half: only the ones the first half doesn't give away. A model that never saw the question can only guess them. A model that memorised it writes the original numbers, which are wrong for the twin.
90
+ - **min-k**: [Min-K% Prob](https://arxiv.org/abs/2310.16789). Average the model's log-probability over the 20% of words that surprise it most. Unseen text always contains some surprising words; memorised text doesn't.
91
+
92
+ ### Calibrating against a model that can't have seen it
93
+
94
+ Even clean models slightly prefer the original numbers, because numbers a person chose fit together better than swapped ones. `contamcheck` cancels this out by subtracting the same preference measured with a **reference model**: the GPT-2 model closest in size (124M to 1.5B parameters), all released in 2019, before GSM8K, HumanEval and MBPP existed. What remains is the familiarity specific to the model under test. The size matters: larger models are better at noticing whether numbers fit together, so a clean 1.4B model looked contaminated against the 124M GPT-2 (20% flagged) but not against the 1.5B one (10%).
95
+
96
+ ## Does it work?
97
+
98
+ A detector that never flags anything would pass every clean model, so `contamcheck` is tested in both directions: clean models must pass and contaminated models must be caught.
99
+
100
+ - **Clean models**: GPT-2, GPT-Neo-125M and Pythia-1.4B were all trained on data collected before GSM8K was released in October 2021, so they cannot have seen it.
101
+ - **Contaminated models**: [`experiments/contaminate.py`](experiments/contaminate.py) takes GPT-Neo-125M and fine-tunes it on 100 of the 200 test questions, for 1, 3 or 10 passes.
102
+
103
+ 200 GSM8K test questions per model, with all defaults (`contamcheck <model> gsm8k`):
104
+
105
+ | Model | Saw the questions? | completion | min-k | Verdict |
106
+ |---|---|---|---|---|
107
+ | GPT-2 (124M) | no | 5% | 0%¹ | ✓ clean |
108
+ | GPT-Neo (125M) | no | 3% | 6% | ✓ clean |
109
+ | Pythia (1.4B) | no | 5% | 10% | ✓ clean |
110
+ | GPT-Neo, 100 leaked × 1 pass | yes | 5% | **18%** | ✗ caught |
111
+ | GPT-Neo, 100 leaked × 3 passes | yes | 8% | **34%** | ✗ caught |
112
+ | GPT-Neo, 100 leaked × 10 passes | yes | **45%** | **53%** | ✗ caught |
113
+ | Qwen2.5 (0.5B) | unknown | 5% | **16%** | ✗ flagged |
114
+
115
+ Percentages are the share of questions flagged. About 5% is expected by chance; **bold** is significant at p < 0.001 (sign-flip permutation test). Completion uses the 154 of 200 questions whose second half contains new numbers. ¹ GPT-2 is its own reference, so min-k can't flag it.
116
+
117
+ All clean models pass, and all contaminated models are caught. The two tests complement each other:
118
+ - **min-k** is sensitive, catching even a single exposure, but relies on the reference model to cancel out other effects.
119
+ - **completion** needs heavier memorisation, but when it fires the evidence is direct. The 10-pass model recalled the exact numbers for 86% of its leaked questions and 5% of the rest.
120
+
121
+ **Qwen2.5-0.5B** is flagged more often by min-k (16%) than any clean model (at most 10%). It doesn't recall numbers verbatim, though. That's consistent with exposure to GSM8K during training, but it isn't proof: the margin over clean models is small, and Qwen was trained on far more maths data than the reference.
122
+
123
+ Contamination from a single exposure is hard to detect. That matches the research: [Duan et al., 2024](https://arxiv.org/abs/2402.07841) find membership inference on LLMs is close to chance when text is seen once. Repeated exposure, which is what happens when a benchmark is copied across many web pages, is caught reliably.
124
+
125
+ ## What I learned building it
126
+
127
+ Every version was checked against clean models (which must pass) and deliberately contaminated ones (which must be caught). Each check found a problem:
128
+
129
+ 1. **The first version flagged GPT-2**, which can't have seen GSM8K. Random replacement numbers read less naturally than numbers a person chose (`8471` vs `150`), so every model preferred the originals. *Fix:* draw replacements from numbers used elsewhere in the benchmark, as often as they're used there, and subtract a reference model's preference.
130
+ 2. **The second version missed models I had contaminated myself.** Training on a question also made the model more familiar with its twin, because they share 90% of their text, so comparing against "all controls" hid the signal. *Fix:* compare each question with *its own* twin and look only at the difference.
131
+ 3. **Word-for-word completion barely noticed a model that had memorised its questions.** It writes nearly the same sentence for the twin too. *Fix:* score only the numbers the first half doesn't give away.
132
+ 4. **A bigger clean model (Pythia-1.4B) was flagged** against the small GPT-2 reference. *Fix:* pick a reference of matching size.
133
+
134
+ ## Limitations
135
+
136
+ - Twins need numbers. For benchmarks without them, pass `--control` with questions the model can't have seen (for example, written after its training cutoff).
137
+ - min-k depends on the reference model cancelling out everything except memorisation. The largest clean reference here is GPT-2 XL (1.5B), so for bigger models pass `--reference` with a larger model trained before the benchmark existed (e.g. `EleutherAI/pythia-6.9b`, trained on 2020 data). Treat a min-k flag on its own as evidence, not proof.
138
+ - Validated so far on GSM8K and models up to 1.5B parameters.
139
+ - Only open-weight models for now. API models don't expose token probabilities for the prompt.
140
+
141
+ ## Development
142
+
143
+ ```bash
144
+ pip install -e ".[dev]"
145
+ pytest
146
+ ```
147
+
148
+ The tests use a fake "cheating" model, so they run in under a second without downloading anything.
149
+
150
+ ## License
151
+
152
+ MIT
@@ -0,0 +1,125 @@
1
+ # contamcheck
2
+
3
+ **Has your language model memorised the benchmark?**
4
+
5
+ When a model scores 90% on GSM8K, the score only means something if the model hasn't seen the test questions before. Benchmarks are public on GitHub, in papers and in blog posts, so they leak into the web-scale data models are trained on. A model that memorised the answers looks smart without being smart.
6
+
7
+ `contamcheck` tests a model for signs that it has seen a benchmark's questions during training.
8
+
9
+ ```
10
+ $ contamcheck experiments/models/neo-gsm8k-10ep gsm8k
11
+
12
+ contamcheck model: experiments/models/neo-gsm8k-10ep benchmark: gsm8k reference: openai-community/gpt2
13
+ 200 questions, control: number twins (23 skipped)
14
+
15
+ ✗ completion: 45% of questions flagged (70/154) — 5% expected by chance (p < 0.001)
16
+ Model writes the exact numbers from the second half of the question
17
+ ✗ min-k: 53% of questions flagged (106/200) — 5% expected by chance (p < 0.001)
18
+ Even the 20% most surprising words aren't surprising to the model (vs reference model)
19
+
20
+ Likely contaminated — the model has probably seen these questions in training
21
+ ```
22
+
23
+ *(A model deliberately trained on half of the GSM8K questions being tested. See [Does it work?](#does-it-work))*
24
+
25
+ ## Install
26
+
27
+ ```bash
28
+ pip install contamcheck
29
+ ```
30
+
31
+ Runs on Apple Silicon (MPS), NVIDIA GPUs (CUDA) or CPU. Any causal language model on the Hugging Face Hub, or a local folder, works.
32
+
33
+ ## Usage
34
+
35
+ ```bash
36
+ contamcheck Qwen/Qwen2.5-0.5B gsm8k # built-in benchmark
37
+ contamcheck Qwen/Qwen2.5-0.5B gsm8k -n 500 # more questions, more power
38
+ contamcheck my-model/ questions.jsonl --field question # your own model and benchmark
39
+ contamcheck my-model/ hf:openai/gsm8k:main:test:question # any Hugging Face dataset
40
+ contamcheck my-model/ humaneval --control fresh.jsonl # control questions written after the model's cutoff
41
+ contamcheck my-model/ gsm8k --json # machine-readable output
42
+ contamcheck big-model/ gsm8k --reference EleutherAI/pythia-6.9b # choose the reference yourself
43
+ ```
44
+
45
+ The exit code is 1 when contamination is likely, so it can gate a CI pipeline.
46
+
47
+ ## How it works
48
+
49
+ ### The twin trick
50
+
51
+ Each benchmark question gets a **twin**: the same text with its numbers swapped for other numbers people used elsewhere in the benchmark.
52
+
53
+ > **Original:** Janet's ducks lay **16** eggs per day. She eats **3** for breakfast...
54
+ > **Twin:** Janet's ducks lay **24** eggs per day. She eats **5** for breakfast...
55
+
56
+ A model that never saw the benchmark has no reason to prefer either version. A model that memorised it prefers the original, because those are the numbers it saw. Each test measures that preference for every question.
57
+
58
+ For a clean model, the preferences land on both sides of zero, and the negative side shows what chance looks like on the positive side. A question is flagged when its preference is stronger than 95% of that mirror image. About 5% of questions get flagged by chance. Far more than that, confirmed by a sign-flip permutation test, means the model has seen the benchmark.
59
+
60
+ ### The tests
61
+
62
+ - **completion**: give the model the first half of the question and check whether it writes the **numbers** of the second half: only the ones the first half doesn't give away. A model that never saw the question can only guess them. A model that memorised it writes the original numbers, which are wrong for the twin.
63
+ - **min-k**: [Min-K% Prob](https://arxiv.org/abs/2310.16789). Average the model's log-probability over the 20% of words that surprise it most. Unseen text always contains some surprising words; memorised text doesn't.
64
+
65
+ ### Calibrating against a model that can't have seen it
66
+
67
+ Even clean models slightly prefer the original numbers, because numbers a person chose fit together better than swapped ones. `contamcheck` cancels this out by subtracting the same preference measured with a **reference model**: the GPT-2 model closest in size (124M to 1.5B parameters), all released in 2019, before GSM8K, HumanEval and MBPP existed. What remains is the familiarity specific to the model under test. The size matters: larger models are better at noticing whether numbers fit together, so a clean 1.4B model looked contaminated against the 124M GPT-2 (20% flagged) but not against the 1.5B one (10%).
68
+
69
+ ## Does it work?
70
+
71
+ A detector that never flags anything would pass every clean model, so `contamcheck` is tested in both directions: clean models must pass and contaminated models must be caught.
72
+
73
+ - **Clean models**: GPT-2, GPT-Neo-125M and Pythia-1.4B were all trained on data collected before GSM8K was released in October 2021, so they cannot have seen it.
74
+ - **Contaminated models**: [`experiments/contaminate.py`](experiments/contaminate.py) takes GPT-Neo-125M and fine-tunes it on 100 of the 200 test questions, for 1, 3 or 10 passes.
75
+
76
+ 200 GSM8K test questions per model, with all defaults (`contamcheck <model> gsm8k`):
77
+
78
+ | Model | Saw the questions? | completion | min-k | Verdict |
79
+ |---|---|---|---|---|
80
+ | GPT-2 (124M) | no | 5% | 0%¹ | ✓ clean |
81
+ | GPT-Neo (125M) | no | 3% | 6% | ✓ clean |
82
+ | Pythia (1.4B) | no | 5% | 10% | ✓ clean |
83
+ | GPT-Neo, 100 leaked × 1 pass | yes | 5% | **18%** | ✗ caught |
84
+ | GPT-Neo, 100 leaked × 3 passes | yes | 8% | **34%** | ✗ caught |
85
+ | GPT-Neo, 100 leaked × 10 passes | yes | **45%** | **53%** | ✗ caught |
86
+ | Qwen2.5 (0.5B) | unknown | 5% | **16%** | ✗ flagged |
87
+
88
+ Percentages are the share of questions flagged. About 5% is expected by chance; **bold** is significant at p < 0.001 (sign-flip permutation test). Completion uses the 154 of 200 questions whose second half contains new numbers. ¹ GPT-2 is its own reference, so min-k can't flag it.
89
+
90
+ All clean models pass, and all contaminated models are caught. The two tests complement each other:
91
+ - **min-k** is sensitive, catching even a single exposure, but relies on the reference model to cancel out other effects.
92
+ - **completion** needs heavier memorisation, but when it fires the evidence is direct. The 10-pass model recalled the exact numbers for 86% of its leaked questions and 5% of the rest.
93
+
94
+ **Qwen2.5-0.5B** is flagged more often by min-k (16%) than any clean model (at most 10%). It doesn't recall numbers verbatim, though. That's consistent with exposure to GSM8K during training, but it isn't proof: the margin over clean models is small, and Qwen was trained on far more maths data than the reference.
95
+
96
+ Contamination from a single exposure is hard to detect. That matches the research: [Duan et al., 2024](https://arxiv.org/abs/2402.07841) find membership inference on LLMs is close to chance when text is seen once. Repeated exposure, which is what happens when a benchmark is copied across many web pages, is caught reliably.
97
+
98
+ ## What I learned building it
99
+
100
+ Every version was checked against clean models (which must pass) and deliberately contaminated ones (which must be caught). Each check found a problem:
101
+
102
+ 1. **The first version flagged GPT-2**, which can't have seen GSM8K. Random replacement numbers read less naturally than numbers a person chose (`8471` vs `150`), so every model preferred the originals. *Fix:* draw replacements from numbers used elsewhere in the benchmark, as often as they're used there, and subtract a reference model's preference.
103
+ 2. **The second version missed models I had contaminated myself.** Training on a question also made the model more familiar with its twin, because they share 90% of their text, so comparing against "all controls" hid the signal. *Fix:* compare each question with *its own* twin and look only at the difference.
104
+ 3. **Word-for-word completion barely noticed a model that had memorised its questions.** It writes nearly the same sentence for the twin too. *Fix:* score only the numbers the first half doesn't give away.
105
+ 4. **A bigger clean model (Pythia-1.4B) was flagged** against the small GPT-2 reference. *Fix:* pick a reference of matching size.
106
+
107
+ ## Limitations
108
+
109
+ - Twins need numbers. For benchmarks without them, pass `--control` with questions the model can't have seen (for example, written after its training cutoff).
110
+ - min-k depends on the reference model cancelling out everything except memorisation. The largest clean reference here is GPT-2 XL (1.5B), so for bigger models pass `--reference` with a larger model trained before the benchmark existed (e.g. `EleutherAI/pythia-6.9b`, trained on 2020 data). Treat a min-k flag on its own as evidence, not proof.
111
+ - Validated so far on GSM8K and models up to 1.5B parameters.
112
+ - Only open-weight models for now. API models don't expose token probabilities for the prompt.
113
+
114
+ ## Development
115
+
116
+ ```bash
117
+ pip install -e ".[dev]"
118
+ pytest
119
+ ```
120
+
121
+ The tests use a fake "cheating" model, so they run in under a second without downloading anything.
122
+
123
+ ## License
124
+
125
+ MIT
@@ -0,0 +1,3 @@
1
+ """contamcheck: check whether a language model has memorised a benchmark."""
2
+
3
+ __version__ = "0.1.0"
@@ -0,0 +1,5 @@
1
+ import sys
2
+
3
+ from .cli import main
4
+
5
+ sys.exit(main())
@@ -0,0 +1,72 @@
1
+ """Load benchmark questions from Hugging Face or from a local file."""
2
+
3
+ from __future__ import annotations
4
+
5
+ import csv
6
+ import json
7
+ from pathlib import Path
8
+
9
+ # name -> (dataset, config, split, field)
10
+ BUILTIN = {
11
+ "gsm8k": ("openai/gsm8k", "main", "test", "question"),
12
+ "humaneval": ("openai/openai_humaneval", None, "test", "prompt"),
13
+ "mbpp": ("google-research-datasets/mbpp", "full", "test", "text"),
14
+ }
15
+
16
+
17
+ def load(spec: str, field: str | None = None) -> list[str]:
18
+ """Load questions from a built-in name, a local .jsonl/.csv/.txt file, or
19
+ `hf:dataset[:config]:split:field` for any Hugging Face dataset."""
20
+ path = Path(spec)
21
+ if path.exists():
22
+ return _load_file(path, field)
23
+ if spec in BUILTIN:
24
+ dataset, config, split, default_field = BUILTIN[spec]
25
+ return _load_hf(dataset, config, split, field or default_field)
26
+ if spec.startswith("hf:"):
27
+ parts = spec[3:].split(":")
28
+ if len(parts) == 3:
29
+ dataset, split, hf_field = parts
30
+ config = None
31
+ elif len(parts) == 4:
32
+ dataset, config, split, hf_field = parts
33
+ else:
34
+ raise ValueError("Use hf:dataset[:config]:split:field, e.g. hf:openai/gsm8k:main:test:question")
35
+ return _load_hf(dataset, config, split, hf_field)
36
+ raise ValueError(f"Unknown benchmark '{spec}'. Built-in: {', '.join(BUILTIN)}; "
37
+ "or pass a .jsonl/.csv/.txt file, or hf:dataset[:config]:split:field")
38
+
39
+
40
+ def _load_hf(dataset: str, config: str | None, split: str, field: str) -> list[str]:
41
+ from datasets import load_dataset
42
+
43
+ ds = load_dataset(dataset, config, split=split)
44
+ if field not in ds.column_names:
45
+ raise ValueError(f"Field '{field}' not in {dataset}. Columns: {', '.join(ds.column_names)}")
46
+ return [str(x) for x in ds[field]]
47
+
48
+
49
+ def _load_file(path: Path, field: str | None) -> list[str]:
50
+ suffix = path.suffix.lower()
51
+ if suffix == ".txt":
52
+ return [line.strip() for line in path.read_text().splitlines() if line.strip()]
53
+ if suffix in (".jsonl", ".json"):
54
+ rows = [json.loads(line) for line in path.read_text().splitlines() if line.strip()]
55
+ elif suffix == ".csv":
56
+ with path.open(newline="") as f:
57
+ rows = list(csv.DictReader(f))
58
+ else:
59
+ raise ValueError(f"Unsupported file type '{suffix}' (use .jsonl, .csv or .txt)")
60
+ if not rows:
61
+ return []
62
+ key = field or _guess_field(rows[0])
63
+ if key not in rows[0]:
64
+ raise ValueError(f"Field '{key}' not in {path.name}. Fields: {', '.join(rows[0])}")
65
+ return [str(r[key]) for r in rows]
66
+
67
+
68
+ def _guess_field(row: dict) -> str:
69
+ for key in ("question", "text", "prompt", "problem"):
70
+ if key in row:
71
+ return key
72
+ raise ValueError(f"Can't tell which field holds the question. Pass --field (one of: {', '.join(row)})")
@@ -0,0 +1,277 @@
1
+ """The contamination tests.
2
+
3
+ Every test gives each question a score where higher means "the model finds this
4
+ more familiar". A score on its own means nothing — models find some text familiar
5
+ just because it is ordinary English — so each benchmark question is compared with
6
+ a control the model cannot have seen.
7
+
8
+ By default the control is the question's "twin": the same text with different
9
+ numbers. The score is then how much more the model likes the original than its
10
+ twin. A model that never saw the benchmark has no reason to prefer either, so its
11
+ scores land on both sides of zero, and the negative side shows what chance looks
12
+ like on the positive side. A question is flagged when its score is higher than 95%
13
+ of the mirrored scores; about 5% get flagged by chance, far more is memorisation.
14
+
15
+ With a control file instead (questions written after the model's training cutoff)
16
+ there are no twins, and a question is flagged if it beats 95% of the controls.
17
+ """
18
+
19
+ from __future__ import annotations
20
+
21
+ import re
22
+ from dataclasses import dataclass, field
23
+ from typing import Callable, Iterator, TypeVar
24
+
25
+ import numpy as np
26
+
27
+ from .model import LanguageModel
28
+ from .perturb import NUMBER
29
+
30
+ PASS, WARN, FAIL = "pass", "warn", "fail"
31
+ EXPECTED_RATE = 0.05 # share of clean questions flagged by chance
32
+ MAX_SUFFIX_WORDS = 40
33
+ MIN_WORDS = 8
34
+
35
+ T = TypeVar("T")
36
+ R = TypeVar("R")
37
+
38
+
39
+ @dataclass
40
+ class Example:
41
+ question: str
42
+ score: float
43
+ expected: str = "" # completion test: the real second half
44
+ generated: str = "" # completion test: what the model wrote
45
+
46
+
47
+ @dataclass
48
+ class TestResult:
49
+ name: str
50
+ description: str
51
+ n: int
52
+ flagged: int
53
+ threshold: float
54
+ p_value: float
55
+ bench_mean: float
56
+ control_mean: float
57
+ examples: list[Example] = field(default_factory=list)
58
+
59
+ @property
60
+ def rate(self) -> float:
61
+ return self.flagged / self.n if self.n else 0.0
62
+
63
+ @property
64
+ def status(self) -> str:
65
+ if self.p_value < 0.001 and self.rate >= 2 * EXPECTED_RATE:
66
+ return FAIL
67
+ if self.p_value < 0.05:
68
+ return WARN
69
+ return PASS
70
+
71
+
72
+ def _threshold(bench: np.ndarray, control: np.ndarray | None) -> float:
73
+ null = -bench if control is None else control
74
+ return float(np.quantile(null, 1 - EXPECTED_RATE))
75
+
76
+
77
+ def _flag_count(bench: np.ndarray, control: np.ndarray | None) -> int:
78
+ return int((bench > _threshold(bench, control)).sum())
79
+
80
+
81
+ def sign_flip_p(diffs: np.ndarray, rounds: int = 2000, seed: int = 0) -> float:
82
+ """Paired test: if the model has no preference between original and twin, each
83
+ difference is as likely to be negative as positive. Flip signs at random and count
84
+ how often the result flags at least as many questions as we observed."""
85
+ observed = _flag_count(diffs, None)
86
+ rng = np.random.default_rng(seed)
87
+ hits = sum(_flag_count(diffs * rng.choice([-1, 1], len(diffs)), None) >= observed for _ in range(rounds))
88
+ return (hits + 1) / (rounds + 1)
89
+
90
+
91
+ def permutation_p(bench: np.ndarray, control: np.ndarray, rounds: int = 2000, seed: int = 0) -> float:
92
+ """Unpaired test: chance of flagging at least as many questions if benchmark and
93
+ control were interchangeable. Shuffling the labels also accounts for the threshold
94
+ being estimated from a finite control set."""
95
+ observed = _flag_count(bench, control)
96
+ pooled = np.concatenate([bench, control])
97
+ rng = np.random.default_rng(seed)
98
+ hits = 0
99
+ for _ in range(rounds):
100
+ rng.shuffle(pooled)
101
+ hits += _flag_count(pooled[:len(bench)], pooled[len(bench):]) >= observed
102
+ return (hits + 1) / (rounds + 1)
103
+
104
+
105
+ def compare(name: str, description: str, bench: np.ndarray, control: np.ndarray, paired: bool) -> TestResult:
106
+ """Flag benchmark questions that look more familiar than chance allows.
107
+ Paired: `bench - control` per question, tested against its own mirror image."""
108
+ if paired:
109
+ diffs = bench - control
110
+ threshold, flagged, p = _threshold(diffs, None), _flag_count(diffs, None), sign_flip_p(diffs)
111
+ else:
112
+ threshold, flagged, p = _threshold(bench, control), _flag_count(bench, control), permutation_p(bench, control)
113
+ return TestResult(
114
+ name=name, description=description, n=len(bench), flagged=flagged, threshold=threshold,
115
+ p_value=p, bench_mean=float(bench.mean()), control_mean=float(control.mean()),
116
+ )
117
+
118
+
119
+ def batched(fn: Callable[[list[T]], list[R]], items: list[T], batch_size: int, desc: str) -> list[R]:
120
+ from tqdm import tqdm
121
+
122
+ out: list[R] = []
123
+ with tqdm(total=len(items), desc=desc, unit="q", leave=False) as bar:
124
+ for batch in _chunks(items, batch_size):
125
+ out.extend(fn(batch))
126
+ bar.update(len(batch))
127
+ return out
128
+
129
+
130
+ def _chunks(items: list[T], size: int) -> Iterator[list[T]]:
131
+ for i in range(0, len(items), size):
132
+ yield items[i:i + size]
133
+
134
+
135
+ # --- Completion test ---------------------------------------------------------
136
+
137
+ def split_question(text: str, max_suffix_words: int = MAX_SUFFIX_WORDS) -> tuple[str, str] | None:
138
+ """Cut a question in half at a word boundary. Keeps the original whitespace so
139
+ code (HumanEval, MBPP) still has its newlines and indentation."""
140
+ words = list(re.finditer(r"\S+", text))
141
+ if len(words) < MIN_WORDS:
142
+ return None
143
+ cut = len(words) // 2
144
+ end = words[min(len(words), cut + max_suffix_words) - 1].end()
145
+ return text[:words[cut].start()].rstrip(), text[words[cut].start():end]
146
+
147
+
148
+ def new_numbers(prefix: str, suffix: str) -> list[str]:
149
+ """Numbers in the second half that don't appear in the first half, so a model
150
+ can only write them by guessing or by remembering."""
151
+ seen = set(NUMBER.findall(prefix))
152
+ return [x for x in NUMBER.findall(suffix) if x not in seen]
153
+
154
+
155
+ def completion_test(model: LanguageModel, bench: list[str], control: list[str], batch_size: int,
156
+ paired: bool = True) -> TestResult:
157
+ """Give the model the first half and check whether it writes the *numbers* of the second half.
158
+
159
+ Comparing whole sentences doesn't work with twins: a model that memorised the
160
+ original writes nearly the same words for its twin, since they share 90% of their
161
+ text. The numbers are what differ. A model that never saw the question can only
162
+ guess them, and is as likely to guess the twin's as the original's; a model that
163
+ memorised it writes the original numbers, which are wrong for the twin."""
164
+ def recall(questions: list[str], desc: str) -> tuple[np.ndarray, list[tuple[str, str, str]]]:
165
+ pairs = [split_question(q) for q in questions]
166
+ generated = batched(lambda b: model.complete(b, 2 * MAX_SUFFIX_WORDS), [p for p, _ in pairs], batch_size, desc)
167
+ scores = []
168
+ for (prefix, suffix), gen in zip(pairs, generated):
169
+ wanted = new_numbers(prefix, suffix)
170
+ written = set(NUMBER.findall(gen))
171
+ scores.append(np.mean([x in written for x in wanted]) if wanted else np.nan)
172
+ return np.array(scores), [(p, s, g) for (p, s), g in zip(pairs, generated)]
173
+
174
+ bench_scores, bench_runs = recall(bench, "completion: benchmark")
175
+ control_scores, _ = recall(control, "completion: control")
176
+ # Questions whose second half brings no new numbers can't be scored
177
+ if paired:
178
+ usable = ~np.isnan(bench_scores) & ~np.isnan(control_scores)
179
+ b, c = bench_scores[usable], control_scores[usable]
180
+ else:
181
+ usable = ~np.isnan(bench_scores)
182
+ b, c = bench_scores[usable], control_scores[~np.isnan(control_scores)]
183
+ result = compare(
184
+ "completion",
185
+ "Model writes the exact numbers from the second half of the question",
186
+ b, c, paired,
187
+ )
188
+ ranked = sorted(np.flatnonzero(usable), key=lambda i: -bench_scores[i])
189
+ for i in ranked[:3]:
190
+ prefix, expected, generated = bench_runs[i]
191
+ if bench_scores[i] > 0:
192
+ result.examples.append(Example(prefix, float(bench_scores[i]), expected, generated))
193
+ return result
194
+
195
+
196
+ # --- Min-K% Prob test ----------------------------------------------------------
197
+
198
+ def min_k_score(logprobs: np.ndarray, k: float = 0.2) -> float:
199
+ """Average log-probability of the k% least likely tokens (Shi et al., 2023).
200
+
201
+ Unseen text always contains some tokens that surprise the model (a name, an odd
202
+ number). Memorised text doesn't: even its rarest tokens get high probability.
203
+ Averaging only the most surprising tokens makes that difference stand out."""
204
+ if len(logprobs) == 0:
205
+ return 0.0
206
+ n = max(1, int(len(logprobs) * k))
207
+ return float(np.sort(logprobs)[:n].mean())
208
+
209
+
210
+ def min_k_test(model: LanguageModel, bench: list[str], control: list[str], batch_size: int,
211
+ paired: bool = True, reference: LanguageModel | None = None, k: float = 0.2) -> TestResult:
212
+ """Min-K% Prob, calibrated against a reference model when one is given.
213
+
214
+ Even with natural replacement numbers, every model slightly prefers the original
215
+ question (its numbers fit together better). Subtracting the score of a reference
216
+ model that provably never saw the benchmark (GPT-2 predates them) cancels that
217
+ out and leaves only the familiarity specific to the tested model."""
218
+ def scores(m: LanguageModel, questions: list[str], desc: str) -> np.ndarray:
219
+ logprobs = batched(m.token_logprobs, questions, batch_size, desc)
220
+ return np.array([min_k_score(lp, k) for lp in logprobs])
221
+
222
+ bench_scores = scores(model, bench, "min-k%: benchmark")
223
+ control_scores = scores(model, control, "min-k%: control")
224
+ if reference is not None:
225
+ bench_scores = bench_scores - scores(reference, bench, "min-k%: reference, benchmark")
226
+ control_scores = control_scores - scores(reference, control, "min-k%: reference, control")
227
+ result = compare(
228
+ "min-k",
229
+ f"Even the {k:.0%} most surprising words aren't surprising to the model"
230
+ + (" (vs reference model)" if reference is not None else ""),
231
+ bench_scores, control_scores, paired,
232
+ )
233
+ ranking = bench_scores - control_scores if paired else bench_scores
234
+ for i in np.argsort(-ranking)[:3]:
235
+ result.examples.append(Example(bench[i], float(bench_scores[i])))
236
+ return result
237
+
238
+
239
+ TESTS = {"completion": completion_test, "min-k": min_k_test}
240
+
241
+
242
+ # --- Building benchmark and control sets ---------------------------------------
243
+
244
+ @dataclass
245
+ class Sets:
246
+ bench: list[str]
247
+ control: list[str]
248
+ control_kind: str # "number twins" or "your control file"
249
+ skipped: int # benchmark questions left out (too short, or no numbers to swap)
250
+
251
+ @property
252
+ def paired(self) -> bool:
253
+ return self.control_kind == "number twins"
254
+
255
+
256
+ def build_sets(questions: list[str], control_questions: list[str] | None, n: int, seed: int = 0) -> Sets:
257
+ """Sample up to n benchmark questions and a matching control set.
258
+
259
+ Without a control file, each control is the question's twin: the same text with
260
+ its numbers swapped for others from the benchmark. Questions without numbers
261
+ have no twin and are skipped."""
262
+ import random
263
+
264
+ from .perturb import NumberPool, perturb
265
+
266
+ rng = random.Random(seed)
267
+ usable = [q for q in questions if split_question(q)]
268
+ if control_questions is None:
269
+ pool = NumberPool(questions)
270
+ pairs = [(q, perturb(q, pool, seed)) for q in usable]
271
+ pairs = [(q, c) for q, c in pairs if q != c]
272
+ sample = rng.sample(pairs, min(n, len(pairs)))
273
+ return Sets([q for q, _ in sample], [c for _, c in sample], "number twins",
274
+ len(questions) - len(pairs))
275
+ controls = [q for q in control_questions if split_question(q)]
276
+ return Sets(rng.sample(usable, min(n, len(usable))), rng.sample(controls, min(n, len(controls))),
277
+ "your control file", len(questions) - len(usable))