Benchmark Lab

Research and methodology from the team

We write about what we find: benchmark design decisions, scoring methodology, and data from our independent evaluation runs. No sponsored content.

All Articles

Public Benchmark API Launch
Product

We Just Opened the Benchmark API to Everyone

Starting today, the Intelligence.AI benchmark data API is publicly accessible on the Free tier. What that means for how teams can integrate scoring into their model selection pipelines.

Priya Natarajan
Code Evaluation Pipeline Walkthrough
Methodology

Inside the Code Evaluation Pipeline: How We Run 400 Tasks Per Model

A walkthrough of the execution infrastructure we built to run code generation benchmarks at scale: task sampling, sandboxed execution, and pass-rate aggregation.

Priya Natarajan
Reasoning Models Benchmark Design
Methodology

Designing Benchmarks for Reasoning Models: What Standard Task Sets Miss

Chain-of-thought models expose weaknesses in benchmarks designed before extended reasoning became mainstream. We rewrote our task structure for the new generation.

Tom Hewson
Image Generation Q2 2026 Results
Research

Image Generation Q2 2026 Results: Nine Models, Three Dimensions of Quality

Quarterly update on image generation rankings. We ran 2,400 prompts across nine models and measured perceptual quality, diversity, and prompt adherence independently.

Tom Hewson
Confidence Intervals Model Evaluation
Methodology

Why We Report Confidence Intervals and Not Just Single Scores

A single benchmark score without a confidence interval tells you almost nothing about reliability. We explain how we compute uncertainty bounds and what they mean for model selection decisions.

Tom Hewson
Video Coherence Scoring Methodology
Methodology

How We Score Video Coherence: Temporal Consistency and Motion Quality Explained

Video generation models need a different scoring framework than image models. We describe the three-axis methodology behind our video coherence leaderboard.

Tom Hewson
Benchmark Drift and Model Updates
Research

Benchmark Drift: What Happens to a Model Score After a Patch Update

We re-ran seven models after their providers released silent patch updates. In four cases, scores shifted by more than three points without any public announcement.

Marius Vollberg
The Case for Independent Benchmarks
Opinion

The Case for Independent Benchmarks in AI Model Selection

When the organization measuring model performance is the same one selling the model, there is a conflict. We wrote this to explain why we think third-party evaluation matters.

Marius Vollberg
Training Data Contamination in Image Benchmarks
Methodology

Training Data Contamination in Image Benchmarks: A Practical Detection Approach

Standard benchmark test sets for image generation can overlap with training data. We describe the contamination detection heuristics we use to qualify our test prompts.

Tom Hewson
Code Correctness vs Code Quality
Methodology

Code Correctness vs. Code Quality: Why Pass Rate Alone Is Not Enough

A model can pass functional tests while generating brittle, unreadable code. We are building a complementary quality dimension into our code benchmark and here is how.

Priya Natarajan
The Reproducibility Problem in ML Evaluation
Research

The Reproducibility Problem in ML Evaluation: What We Learned from Rerunning Published Benchmarks

We attempted to reproduce 12 published benchmark results using original methodology descriptions. Eight of twelve required undocumented assumptions to replicate.

Tom Hewson
Task Coverage Map for Code Benchmarks
Research

A Task Coverage Map for Code Benchmarks: What Problem Types Are Actually Tested

Not all code generation benchmarks test the same things. We built a taxonomy of task types and mapped coverage across six major benchmarks to show where the gaps are.

Tom Hewson

New articles every few weeks

Benchmark results, methodology explainers, and research findings. Notify me when something new publishes.