Benchmark Lab Research

Image Generation Q2 2026 Results: Nine Models, Three Dimensions of Quality

Tom Hewson  · 

Abstract grid of varying image quality scores represented as color-coded tiles

Quarterly update on image generation rankings. We ran 2,400 prompts across nine models and measured perceptual quality, diversity, and prompt adherence independently.

Q2 2026 was the first evaluation cycle where we ran our three-dimensional image scoring at full scale: 2,400 prompts, nine models, three independent quality axes measured separately rather than collapsed into a single composite score at evaluation time. The results confirmed something we had been seeing in partial-run data: overall rankings depend heavily on which quality dimension you weight most, and the dimension that matters most varies by use case.

The three axes explained

We score image quality on perceptual quality, diversity, and prompt adherence. These are genuinely independent measures, not three ways of measuring the same thing.

Perceptual quality is the closest to what people mean when they say "does this look good." We use a combination of FID (Frechet Inception Distance) against a held-out reference distribution, a learned perceptual metric, and a frequency-domain coherence measure that catches high-frequency artifacts (commonly visible as noise patterns in hair, fabric, and background textures). FID rewards distributional similarity; the perceptual metric rewards local feature quality; the frequency measure catches specific failure modes that FID tends to miss. We weight these three sub-measures equally and report the composite as the perceptual quality score.

Diversity measures the variation within a model's output across a fixed prompt set. A model that generates the same compositional template for every portrait prompt, with only minor color and feature variation, produces images that might each score highly on perceptual quality but that collectively provide low diversity. Diversity is measured as the average pairwise distance in feature space across 50 outputs for each of 48 prompts, normalized to a 0-100 scale. Low diversity is sometimes a design choice: some models are tuned for consistency. But it affects use cases where creative variation matters.

Prompt adherence measures how closely the output reflects what the prompt requested. We use a suite of structured prompts with explicit, verifiable requirements: "a photograph of a red chair next to a window, afternoon light, no other furniture visible." An automated pipeline checks each output against the required elements: is the chair present? is it red? is there a window? is there other furniture? Each required element either passes or fails, and the adherence score is the proportion of elements that pass across all structured prompts. This is not a perfect measure for creative prompts where "adherence" is more subjective, but for specification-following tasks it is reliable.

Q2 results by model

Nine models completed the full Q2 evaluation. Scores below are from our internal benchmark runs; confidence intervals at 95 percent are based on 2,400-prompt samples.

lumina-7b led on perceptual quality (score: 78, CI: 75-81), maintaining the top position it held in Q1. Its FID against our reference distribution was the lowest in the set (22.1), and its frequency-domain coherence was notably clean. Its diversity score was the lowest of the nine models (42), which we flagged in its leaderboard entry. For use cases where output consistency is more important than creative variation, this is a feature. For generative creative workflows, it is a limitation.

cascade-vision scored highest on diversity (71) and prompt adherence (79), placing it second overall on the composite. It generated more compositionally varied outputs per prompt than any other model, and it was the most reliable at following explicit structural requirements. Its perceptual quality score (63) lagged the top two models, with the frequency-domain coherence measure identifying recurring texture artifacts in synthetic materials.

aurora-pro showed the largest gap between its vendor-reported score and our measurement. We covered this in detail in the vendor benchmark inflation article. On our Q2 run, it placed sixth overall with a composite of 58, despite a vendor-reported FID of 18.4 that would have implied a top-three placement.

The remaining six models clustered in the 50-67 composite range. Three of the six showed significant year-over-year improvement from their Q2 2025 scores (for models present in both cycles); one showed regression, with a drop of 7 composite points that correlates with a patch update we documented in our benchmark drift tracking.

What moved most between Q1 and Q2

The biggest change was in prompt adherence scores across the field. Six of nine models improved their adherence scores by 4 points or more compared to Q1, which suggests that the current generation of image models is being fine-tuned with more structured prompt following in mind. This is a real capability improvement for production use cases.

Perceptual quality scores were more stable. The top three models on perceptual quality in Q2 were the same as in Q1, with rank-order preserved. The underlying feature quality improvements seem to be slower-moving than the prompt adherence gains.

Diversity scores declined slightly across the field. Five of nine models produced lower diversity scores in Q2 than Q1. We do not have a confident explanation for this. One hypothesis is that fine-tuning for prompt adherence tends to reduce output variation because it optimizes for a specific correct answer rather than a range of valid outputs. The adherence-diversity tradeoff is a real tension in image model design, and the Q2 data suggests current model development is resolving it toward adherence.

Methodology notes for this cycle

We added 200 new prompts to the Q2 set compared to Q1. The additions were concentrated in two categories where we felt coverage was thin: text-in-image prompts (where the model must generate readable text as part of the image) and prompts specifying lighting conditions explicitly. Both categories have proven to be reliable differentiators between models. Text legibility in image generation remains inconsistent across the field, with our internal runs showing a range of 18 to 64 percent success rate on readable-text prompts depending on the model.

We also changed our human preference sampling for this cycle. In Q1, we used a panel of five evaluators for the subjective quality assessments that feed into our perceptual metric calibration. In Q2, we expanded to twelve evaluators and stratified the panel by professional background (graphic designers, photographers, and general consumers in roughly equal proportions). The calibration results were meaningfully different across groups, particularly on the diversity measure. We report the aggregate, but we are tracking the per-group calibrations internally to understand how professional and consumer preferences diverge.

One limitation to flag

Our nine-model sample is not a complete picture of the image generation field. We only benchmark models we can access programmatically at our tier of API access, which excludes some models that are available only through proprietary platforms or that require special access agreements. The leaderboard ranks what we can measure, not everything that exists. Teams doing final model selection should treat our ranking as a starting filter, not a final verdict, particularly if they have access to models not in our set.

See the current image rankings

Our image quality leaderboard updates weekly with new scores across all three dimensions.

View Image Leaderboard