Benchmark contamination is the presence in training data of evaluation items, or near-copies of them, so that a score measures memory rather than ability. The public version of the problem, a benchmark leaking into a pretraining crawl, gets the attention. The private version is more common and entirely self-inflicted: a team builds its own evaluation set by randomly splitting rows, and the training side contains near-duplicates of the evaluation side.
This post explains how a row-level split leaks, what the right unit of splitting is in each domain, how to decontaminate against public benchmarks, and what an overlap report should contain so that a reader can trust the number.
How benchmark contamination gets in through a row-level split
A random split assigns each row independently to training or evaluation. That is correct only if rows are independent, and in real corpora they are not: the same file exists in a repository and its forks, the same filing is restated the next quarter, the same case is quoted in later judgments, and the same robot session is cut into episodes that share a scene and an operator. A random split puts members of each of those groups on both sides.
The effect is measurable. Lee et al. (2021), deduplicating standard language-modelling datasets, found train-test overlap affecting over 4% of the validation set, and models trained on the deduplicated data emitted memorised text ten times less often. A validation set with that much overlap rewards memorisation on every item that overlaps, and the overall score moves in the direction of whichever model memorised more.
Exact deduplication before splitting does not fix this, because the leaking records are near-duplicates and semantic duplicates, which exact hashing never sees. The three tiers of deduplication explain what each catches; the point here is that the split has to be designed so that the misses of all three tiers land on the same side.
Split by repository, document or entity
The unit of the split is the thing that generates correlated records, and it differs by domain:
- Code: the repository, extended to its fork family, so that a file and its forked copies are on one side. Evaluation tasks come from repositories that contribute nothing to training.
- Documents: the source document, and often the source itself, so that two chunks of one filing or two articles from one template cannot straddle the split.
- Entities: the company, the case, the protein, the patient, so that every record about one entity is on one side. This matters most when the task is to predict something about the entity.
- Time: for anything time-ordered, a cut-off date with evaluation after it, so that the model is not graded on a past whose future it has seen.
The mechanics are simple once the key exists: assign groups, not rows, and keep the group key in each record's provenance so the assignment can be reproduced. The order of operations matters. Deduplicate first, across the whole corpus; then split by group; then run near-duplicate detection across the two sides as a check. The check finds what the group key missed: a file copied between unrelated repositories, a document quoted in another, an entity with two identifiers.
The cost is a less convenient evaluation set. Group splits produce uneven sizes and leave some groups entirely on one side, so the evaluation set covers fewer sources than a random split would. That is the correct trade: an evaluation set that covers every source is one that shares every source with training.
Decontamination against public benchmarks
The second obligation is to public benchmarks: a training set must not contain the items a model will later be scored on. The usual check is n-gram overlap, comparing token sequences from each benchmark item against the training data and removing the matches. It is cheap, and it is known to be insufficient.
Yang et al. (2023) show that simple variations of test data, paraphrasing or translation, bypass n-gram decontamination, and that a 13B model trained on rephrased test sets reached scores on par with GPT-4. They also report that 8% to 18% of HumanEval overlapped with public pretraining corpora they examined. The practical consequence is that decontamination needs the same three tiers as deduplication: exact and n-gram matching for the copies, embedding similarity for the rephrasings, and a human look at the candidates.
For code tasks it also means keeping out of training the repositories that public benchmarks draw on, such as the repositories SWE-bench takes its tasks from; the code domain rules state that evaluation tasks are checked for overlap with public benchmarks. Keep a registry of the benchmarks decontaminated against, with versions, because benchmarks change and new ones appear after a dataset is built. A training set decontaminated against last year's list is contaminated against this year's.
What decontamination still cannot prove
A clean split proves that your evaluation set is independent of your training set. It says nothing about the pretrained model you start from, which may have seen the public benchmark, or a paraphrase of it, long before your data arrived. The honest position is to report both: the overlap between your own splits, and the fact that the base model's exposure is unknown unless its training data is published. Telling whether a model change made things worse depends on an evaluation set whose independence you can actually assert.
How to report overlap
An overlap report is what turns "decontaminated" from an adjective into a claim with evidence. It should state:
- The group key used for the split, per domain, and how group membership was determined.
- The methods run across the split and against each public benchmark: exact, n-gram size, MinHash threshold, embedding model and similarity threshold.
- The benchmarks and versions checked.
- Counts found per method, and the action taken: removed from training, or kept and flagged, with the reason.
- The sample of candidate pairs a person reviewed, and what share were true overlaps.
- What was not checked, and why.
A report that gives a single overlap percentage with no method, no benchmark list and no denominator cannot be checked, and a data quality report that cannot be checked is a brochure.
Runix Data splits evaluation data from training data by source, so that near-duplicates cannot sit on both sides, and, in the code domain, checks evaluation sets for overlap with public benchmarks; the quality report that accompanies every delivery states what was dropped and why. Runix Data is in early access, by engagement, priced per project or by volume and quoted before any work starts.
Questions this raises
Why is a random train and evaluation split a problem?
Rows in real corpora are not independent: forks, restated filings, quoted cases and episodes from one session are near-duplicates of each other. A random split puts members of each group on both sides, so the evaluation score partly measures memorisation.
What should the unit of the split be?
The thing that generates correlated records: the repository for code, the document or source for text, the entity for records about companies, cases or patients, and a cut-off date for anything time-ordered.
Is n-gram decontamination against public benchmarks enough?
No. Paraphrased or translated copies of benchmark items pass n-gram checks, so decontamination needs embedding similarity and human review as well, plus a versioned registry of the benchmarks that were checked.