Benchmark Lab Methodology

Why We Report Confidence Intervals and Not Just Single Scores

Tom Hewson  · 

Abstract data visualization showing confidence band ranges around model performance measurements

A single benchmark score without a confidence interval tells you almost nothing about reliability. We explain how we compute uncertainty bounds and what they mean for model selection decisions.

Most published benchmark scores are point estimates. A model scores 76.3 on a given task suite. Another scores 74.8. The naive reading is that the first model is better. But without knowing the variance of those measurements, you cannot tell whether the 1.5-point difference is a real capability gap or statistical noise in task sampling. We report confidence intervals on all our leaderboard entries specifically to avoid that ambiguity. Here is how we compute them and why the width of the interval often matters more than the point estimate.

Why point estimates alone are misleading

A benchmark score is a sample statistic. You ran N tasks, counted how many the model passed, and divided. That number is your estimate of the model's true pass rate on the distribution of tasks that the benchmark is supposed to represent. Like any sample statistic, it has uncertainty. Run the same model on a different random sample of 400 tasks from the same pool, and you will get a different point estimate. How different depends on the variance in the model's per-task performance.

Models with highly variable per-task performance produce wide confidence intervals from the same sample size. A model that is consistently good (passes most tasks reliably) produces a narrow interval. A model that is strong on some task types and weak on others will show a wider interval, because its aggregate score shifts more depending on which tasks happen to be in the sample. Reporting only the point estimate treats these two cases the same way, which misleads about reliability.

There is also a practical selection problem. Suppose you are choosing between nexus-code-4 (score: 77.2, CI: 74.1-80.3) and orion-r1 (score: 76.0, CI: 75.2-76.8). The point estimates favor nexus-code-4, but the confidence intervals overlap substantially. Nexus-code-4's interval extends lower than orion-r1's entire interval. Orion-r1's interval is tighter, indicating more consistent performance. Depending on whether you weight average performance or consistency more, the "better" model could be either one. Without the intervals, you see only the 1.2-point difference and conclude nexus-code-4 is clearly better. That conclusion is not supported by the data.

Our bootstrapping procedure

We use bootstrapped confidence intervals rather than parametric intervals because we do not assume a specific distributional form for per-task outcomes. The bootstrap is non-parametric: it estimates the sampling distribution of our statistic (pass rate) by resampling from the observed data repeatedly.

The procedure for a single model evaluation run is as follows. We start with the binary pass/fail outcomes for all 400 tasks in the run. We then draw 1,000 bootstrap samples, each by sampling 400 tasks with replacement from the original 400. For each bootstrap sample, we compute the pass rate. The distribution of 1,000 bootstrap pass rates forms our estimate of the sampling distribution. The 2.5th and 97.5th percentiles of that distribution are the lower and upper bounds of the 95 percent confidence interval.

1,000 bootstrap samples is sufficient for stable interval estimation at our sample sizes. We ran comparisons at 500, 1,000, 5,000, and 10,000 samples for several of our model evaluations and found that interval bounds stabilize well before 1,000 samples. The additional computation cost of going beyond 1,000 does not produce meaningfully different intervals for our use case.

What the intervals mean for the leaderboard

We annotate leaderboard entries where two models have overlapping confidence intervals with a "statistically indistinguishable" flag. This is not a qualitative judgment; it is a mechanical check. If the lower bound of the higher-scoring model overlaps with the upper bound of the lower-scoring model, we flag both entries. Users who see that flag know the rank ordering is not reliable at the current sample size.

The most common source of overlap in our code leaderboard is between models ranked 3-5 and ranked 6-8. The top two spots are usually separated from the rest by margins that exceed even wide confidence intervals. The bottom of the top tier and the top of the middle tier frequently swap positions across weekly runs, and the confidence interval overlap explains why: there is not a meaningful capability difference between models at ranks 4 and 6, only sampling variation.

This information should affect how you use the leaderboard. If you are choosing between a model in the clearly-top cluster and a model in the statistically-indistinguishable middle cluster, the ranking tells you something real. If you are choosing between two models both in the indistinguishable middle, the leaderboard rank should not be your deciding factor. Use the per-task-category subscores (available via the API) to see if one model has a specific strength in your task family.

Sample size effects

Our standard evaluation run uses 400 tasks. We chose 400 through a calibration analysis of interval width versus compute cost. At 200 tasks, the bootstrapped intervals are roughly 1.5x as wide as at 400 tasks for models with typical performance variance. At 800 tasks, they narrow by roughly 30 percent compared to 400. The computational cost scales linearly with tasks, but the interval improvement is sublinear (width scales as roughly 1/sqrt(N) for a fixed pass rate). We judged that the interval width at 400 tasks is narrow enough to distinguish genuinely different models, and that the cost of doubling to 800 tasks for marginal interval improvement is not justified given our weekly run cadence across 28 models.

For teams that need tighter confidence bounds because they are making high-stakes decisions between closely-ranked models, the Pro tier includes access to our 800-task supplemental runs, which we conduct quarterly rather than weekly. Those runs have roughly 30 percent narrower confidence intervals and are the right input for a selection decision where a 1-point difference in score matters.

A note on interpreting multi-dimensional scores

For image and video models, our confidence intervals cover the composite score. Each underlying dimension (perceptual quality, diversity, prompt adherence for image; temporal consistency, motion quality, frame coherence for video) has its own interval. The composite interval is derived from a weighted combination of the component intervals, and it tends to be narrower than the individual component intervals because some of the per-component variance is uncorrelated and cancels when aggregated.

When you are selecting an image model for a use case that depends heavily on prompt adherence specifically, look at the confidence interval for the adherence sub-score, not the composite. A model might have a narrow composite interval (appears consistent overall) while having a wide adherence interval (variable in exactly the dimension you care about). The full component score breakdown with component-level confidence intervals is accessible through the API for all models we track.

What we do not try to quantify

Confidence intervals address sampling uncertainty: the noise in our score estimate from using a finite task sample. They do not address specification uncertainty: the possibility that our task distribution itself is not representative of the real-world use case you care about. A model might have a tight, reliable confidence interval on our task set and perform completely differently on your actual task distribution. Our benchmark is designed to be broadly representative, but it is not infinitely granular. For mission-critical model selection, our benchmark is a starting point, and internal evaluation on your specific tasks remains the right final step. We are trying to make the starting point more reliable, not replace the ending step.

See scores with confidence intervals

Our leaderboards display uncertainty bounds alongside every score.

View Leaderboards