Methodology

How we run benchmarks

Independent, reproducible evaluations with published scoring criteria. No vendor funding, no cherry-picked test sets, no undisclosed assumptions.

Active evaluations 3
Total models tracked 28
Tasks run to date 140K+
All results include 95% confidence intervals
Principles

What makes a benchmark trustworthy

Every evaluation decision is made against these four principles. They are not aspirational. They are gates.

Vendor independence

No model provider funds or influences any evaluation. We purchase API access at standard rates. Providers have no advance visibility into our test sets before publication.

Published methodology

Every scoring formula, weighting decision, and data pipeline step is documented before a test cycle begins. We do not adjust methodology based on results.

Reproducibility first

We publish task IDs, sampling seeds, and execution parameters so independent researchers can verify our results. If a result cannot be reproduced with the same inputs, we flag it.

Uncertainty reporting

Every score comes with a 95% confidence interval. A model with a 2-point lead and overlapping confidence intervals is not statistically ahead. We state that explicitly.

Regular re-testing

Models are re-evaluated when providers release updates. Silent patch updates that shift scores are tracked and disclosed. We do not rely on a single historical measurement.

Contamination control

Test sets are refreshed quarterly using novel tasks not seen in public datasets. We apply heuristic contamination detection to flag prompts that may have appeared in training data.

Pipeline

How an evaluation runs

From task selection to leaderboard publication, each cycle follows a fixed five-step process.

Task Selection 400-2400 tasks decontaminated Execution Sandboxed API calls fixed parameters Scoring Automated + panel per published formula Statistics 95% CI computed outlier review Publication Leaderboard updated data available via API STEP 1 STEP 2 STEP 3 STEP 4 STEP 5
Scoring Detail

Score composition by evaluation type

Each evaluation has a different composite formula reflecting what actually matters for that output type.

Code Generation

Pass@1 rate 60%
Code quality 25%
Edge-case handling 15%

400 tasks per model. Quality scored by static analysis for readability, type correctness, and anti-pattern absence.

Image Quality

Perceptual quality 50%
Prompt adherence 30%
Output diversity 20%

2,400 prompts per model across four content categories. Panel scoring applied to 200 sampled outputs per cycle.

Video Coherence

Temporal consistency 45%
Motion quality 35%
Prompt fidelity 20%

800 clips per model, 4-10 seconds each. Temporal consistency measured by optical flow deviation across frame pairs.

Get the underlying data

Raw scores, confidence intervals, and per-task breakdowns available via the API. Pro plan includes weekly refresh access; Enterprise adds custom evaluation runs.

See Plans View Leaderboards