How we run benchmarks
Independent, reproducible evaluations with published scoring criteria. No vendor funding, no cherry-picked test sets, no undisclosed assumptions.
What makes a benchmark trustworthy
Every evaluation decision is made against these four principles. They are not aspirational. They are gates.
Vendor independence
No model provider funds or influences any evaluation. We purchase API access at standard rates. Providers have no advance visibility into our test sets before publication.
Published methodology
Every scoring formula, weighting decision, and data pipeline step is documented before a test cycle begins. We do not adjust methodology based on results.
Reproducibility first
We publish task IDs, sampling seeds, and execution parameters so independent researchers can verify our results. If a result cannot be reproduced with the same inputs, we flag it.
Uncertainty reporting
Every score comes with a 95% confidence interval. A model with a 2-point lead and overlapping confidence intervals is not statistically ahead. We state that explicitly.
Regular re-testing
Models are re-evaluated when providers release updates. Silent patch updates that shift scores are tracked and disclosed. We do not rely on a single historical measurement.
Contamination control
Test sets are refreshed quarterly using novel tasks not seen in public datasets. We apply heuristic contamination detection to flag prompts that may have appeared in training data.
How an evaluation runs
From task selection to leaderboard publication, each cycle follows a fixed five-step process.
Score composition by evaluation type
Each evaluation has a different composite formula reflecting what actually matters for that output type.
Code Generation
400 tasks per model. Quality scored by static analysis for readability, type correctness, and anti-pattern absence.
Image Quality
2,400 prompts per model across four content categories. Panel scoring applied to 200 sampled outputs per cycle.
Video Coherence
800 clips per model, 4-10 seconds each. Temporal consistency measured by optical flow deviation across frame pairs.
Get the underlying data
Raw scores, confidence intervals, and per-task breakdowns available via the API. Pro plan includes weekly refresh access; Enterprise adds custom evaluation runs.