Benchmark Lab Methodology

Training Data Contamination in Image Benchmarks: A Practical Detection Approach

Tom Hewson  · 

Abstract visual showing data overlap detection between training sets and evaluation benchmarks

Standard benchmark test sets for image generation can overlap with training data. We describe the contamination detection heuristics we use to qualify our test prompts.

Image generation benchmark contamination is different from language model benchmark contamination in an important way. For language models, contamination means the model has seen the exact question and answer during training. For image generation models, contamination means the training set contains images that are semantically very close to what the benchmark prompts are requesting. The model may not have memorized specific benchmark prompts, but if the reference distribution for a benchmark overlaps heavily with training data, the model is not generalizing when it scores well on that benchmark. It is interpolating.

We designed our image evaluation prompts specifically to be resistant to this kind of overlap. This post describes the three main signals we use to detect contamination in prompt qualification, and what we do when a prompt fails.

Why contamination matters differently for generation vs. classification

In image classification benchmarks, contamination is relatively well-studied: if the test images appeared in training, classification accuracy on those images measures memorization. Detection approaches like membership inference attacks or training data fingerprinting are mature enough to catch most cases.

For generation benchmarks, the problem is more subtle. We are not checking whether a specific image was memorized. We are checking whether the prompt distribution we use to elicit images is semantically similar to what the model was trained on. A model trained on 100 million internet photographs will have very dense coverage of common visual categories: everyday objects, landscapes, human faces, domestic interiors. If our benchmark prompts cluster heavily in those categories, the model's scores on our benchmark reflect its performance in the highest-density regions of its training distribution, not its general generative capability.

The practical consequence is that FID scores in particular can be artificially lowered (recall: lower FID is better for image quality) by models with large training sets that match our reference distribution closely, without that low FID reflecting superior perceptual generation quality. The model looks good because it is close to what it has seen, not because it generates especially well.

Signal 1: Prompt distribution overlap with common web image categories

Our first contamination check is a distributional one. We maintain a taxonomy of visual content categories derived from public dataset documentation for large-scale image training sets. For each candidate benchmark prompt, we classify the prompt into this taxonomy and check its category density. If a prompt falls into a category that is known to be extremely well-represented in typical web-scraped training sets (everyday objects, common domestic scenes, generic landscape photography), we apply a difficulty adjustment or replace the prompt with a variant from a lower-density category.

This does not mean we exclude common categories entirely. It means we ensure our prompt set is distributed across the full category spectrum, not concentrated in the high-density regions where contamination effects are strongest. We target a category distribution that is roughly flat across our taxonomy, which is deliberately different from the category distribution of any known large-scale training set.

Signal 2: Output memorization patterns in high-fidelity prompts

The second signal comes from analyzing model outputs on specific prompts that we have crafted to be highly specific about composition and subject matter. If a prompt describes a scene with enough specificity that there are only a small number of distinct valid interpretations, a model that has memorized training examples closely related to those interpretations will tend to produce outputs that cluster tightly in feature space.

We measure this by running each specific prompt 20 times across the models we evaluate and computing the feature space variance of the 20 outputs. A model with strong creative generalization should produce diverse outputs for a specific prompt because it is drawing on a rich learned distribution. A model that has effectively memorized the relevant training examples tends to produce outputs that cluster, because it is sampling near a high-probability mode in its learned distribution rather than exploring the full valid output space.

We flag prompts where any model in our evaluation set shows anomalously low output variance (more than 2 standard deviations below the mean variance for that model across all prompts). A flagged prompt is reviewed, and if the low variance pattern is consistent across multiple models, we suspect the prompt is in a region of high training density and rotate it out of the evaluation set.

Signal 3: Temporal consistency of scores across prompt cohorts

The third signal is longitudinal. We divide our prompt set into cohorts based on when each prompt was introduced to our evaluation pool. Prompts in our oldest cohort (introduced when we launched image evaluation in early 2025) have been part of the public benchmark longer and are more likely to have influenced subsequent training data updates if providers were training on benchmark-adjacent data. New prompt cohorts introduced more recently should be harder for the same models.

If a model's scores on old-cohort prompts are substantially higher than its scores on new-cohort prompts (controlling for prompt difficulty), that divergence is consistent with the old prompts being more contaminated. We look for per-model cohort divergences greater than 6 points as a flag. Two of the nine models in our Q2 evaluation showed cohort divergences above this threshold. We do not attribute this definitively to contamination, but we track it as a risk indicator and weight new-cohort prompt scores more heavily in our composite calculation for those models.

What we do when a prompt fails qualification

Prompts that fail any of the three signals are quarantined for review. A quarantined prompt is not automatically removed; it is re-evaluated against the specific model outputs that triggered the flag. Some prompts that look contaminated are just prompts for which our models happen to perform particularly consistently, and consistently good is not the same as contaminated. We remove prompts that show strong evidence of training data overlap from our active evaluation set and replace them from our qualification pool.

We rotate roughly 8 percent of our image prompt set each evaluation cycle. Some of that rotation is contamination-driven removal; some is planned diversity cycling. This rotation means our benchmark scores are not perfectly comparable across cycles (a model run in Q1 and Q2 faced slightly different task sets), which is why we compute cross-cycle comparable scores by holding a fixed anchor prompt set that we never rotate and using it to calibrate the score normalization across cycles.

Limitations of our contamination detection

Our three signals are heuristics, not proofs. We cannot definitively determine whether a specific prompt is contaminated without access to a model's training data, which we do not have. What we can do is use multiple indirect signals to identify prompts that are likely to produce inflated scores and rotate them out before they degrade our benchmark's quality. Perfect contamination resistance would require generating entirely novel prompts from distributions that no model has ever trained on, which is practically impossible given the scale of modern training sets. We are managing a probabilistic contamination risk, not eliminating it.

We publish our prompt qualification methodology so that researchers and evaluation teams can assess whether our approach is sufficient for their purposes. If you are using our image benchmark data for research purposes and have specific questions about prompt provenance, contact us directly. We maintain detailed records of when each prompt was introduced and which qualification checks it passed.

Benchmarks built to resist contamination

Our image evaluation sets rotate regularly and include novel prompts not found in public training datasets.

Explore Benchmarks