Benchmark Lab Research

Vendor Benchmark Inflation: Why Self-Reported Scores Keep Missing in Real Tasks

Marius Vollberg  · 

Abstract illustration representing inflated benchmark scores diverging from real-world performance

We tracked 18 model releases where vendor-claimed scores diverged from our independent measurements by more than 8 points. Here is what we found.

The pattern started showing up in mid-2025, and we documented it carefully because we wanted to be sure we weren't misconfiguring our own test harness. After ruling that out, the finding held: there is a consistent gap between the scores vendors put in press releases and the scores those same models produce when we run our standard task suite independently. The median gap in our 18-release sample was 11.3 points. That is not noise.

How the gap develops

We are not saying vendors lie. The more accurate framing is that score selection and benchmark construction create a systematic upward pull that no single dishonest decision causes. Three mechanisms show up repeatedly in our analysis.

First, vendors benchmark on development sets they have had visibility into during training. This is not always deliberate contamination. It can be as subtle as a data curation team removing "low quality" examples that happen to overlap heavily with a public benchmark's harder tasks. The model never sees the exact test items, but the distribution it trained on has been shaped by what the benchmark rewards. Our held-out sets are drawn from sources the models have had no prior exposure to, which is why our numbers tend to be lower.

Second, vendors typically report scores from their best-configured inference run. Temperature, top-p, chain-of-thought prompting, and few-shot examples are all tunable. If you run 40 configurations and report the top result, you have introduced an evaluation-time optimization that end users will not replicate. We run a fixed inference configuration for all models, same prompt format, same generation parameters, applied identically. When we benchmarked nexus-code-4 on function completion tasks, the vendor's reported pass rate was 84.1 percent. Under our fixed-configuration run, it came in at 72.6 percent. Neither number is wrong in an absolute sense, but only one of them predicts what a developer will see when they call the API with default parameters.

Third, task selection effects are real and underappreciated. A benchmark suite covers some problem types well and others barely at all. If a model is particularly strong at, say, algorithm completion problems, and a benchmark happens to weight that category heavily, the reported headline number will outperform the model's actual general utility. We build task coverage maps specifically to detect this (see our separate writeup on task coverage), and we reweight categories to maintain even coverage across problem types.

The aurora-pro case study

The clearest example in our dataset came from aurora-pro, an image generation model with a vendor-reported FID score of 18.4 on a standard perceptual quality benchmark. FID (Frechet Inception Distance) is a lower-is-better metric, and 18.4 would place the model among the top three on our leaderboard at the time.

We ran aurora-pro through our image evaluation pipeline, using 600 held-out prompts stratified across six content categories: natural scenes, synthetic objects, architectural spaces, abstract patterns, human-adjacent subjects, and text-in-image. Our measured FID was 27.1. On the perceptual quality composite score, where we weight visual coherence, artifact frequency, and prompt adherence independently, aurora-pro scored 61 out of 100, landing in the middle tier of our nine-model Q2 ranking.

The vendor's test set, as far as we could reconstruct it from their methodology note, used 200 prompts concentrated in natural scenes and synthetic objects, the two categories where aurora-pro is genuinely strongest. Our stratified sample drew out the model's weaker performance on text-in-image prompts and complex multi-object scenes. Neither of those categories is exotic; they show up constantly in real workflows. The difference between 18.4 FID and 27.1 FID is the difference between "leading model" and "solid mid-tier."

What this is not

We want to be precise about the claim we are and are not making. We are not saying that vendor benchmarks are fabricated or that the models are bad. Several models in our dataset genuinely outperform the field on certain task types, and their vendor-reported numbers accurately reflect that on those specific tasks. The problem is when a number derived from a favorable task subset or favorable inference configuration gets used as a general capability claim.

We are also not saying that independent third-party benchmarks are the only valid form of evaluation. For a team with very specific use cases, a custom internal evaluation against their own task distribution will outperform any general benchmark, ours included. Our leaderboard is useful precisely when you do not have the resources to run 400 custom tasks per model yourself, and when you need a comparison across models that is apples-to-apples.

The structural incentive problem

Benchmark inflation is not primarily a technical problem. It is an incentive problem. Model releases are competitive events, and the press cycle rewards high numbers. A vendor that reports conservative, held-out, fixed-inference scores will look worse than a competitor reporting favorable-configuration numbers, even if the underlying model capabilities are identical. The rational response to that incentive structure is to optimize the benchmark presentation, not the model.

This is why third-party evaluation cannot be optional. When the same organization builds the model, selects the benchmark, chooses the inference configuration, and reports the score, every decision in that chain is made by an entity with a stake in the outcome. The score is not independent evidence. It is a marketing artifact with a rigorous-looking number attached to it.

We have had conversations with evaluation engineers at growing AI teams who describe the same experience: they read a model release, try the model on their actual use case, and find results 10 to 20 points below the announced benchmark. They then spend time figuring out whether their integration is wrong. Usually it isn't. They just hit the gap between benchmark-optimized evaluation and real inference conditions.

What we changed in our methodology after this analysis

Documenting the gap prompted us to tighten two things. First, we now publish our exact inference configuration for every model run, including the version of the model we called, the generation parameters, the prompt template, and the date of evaluation. This makes our numbers reproducible: anyone with API access to the model can re-run our configuration and check our score. Second, we added a "configuration gap" annotation to the leaderboard for models where we know the vendor ran a materially different inference setup. This is not a penalization. It is context. Users can see the vendor's number, our number, and what accounts for the difference.

Neither change is complete. There will always be models we cannot fully characterize, task distributions that we have not covered, and inference configurations we have not explored. But the commitment to fixed, documented, reproducible evaluation conditions is what separates a benchmark from a benchmark-shaped marketing document. That distinction is what we are building here, and the 11-point median gap in our dataset is the reason it matters.

See the independent numbers

View current leaderboard rankings across code, image, and video generation tasks.

View Leaderboards