Leaderboard

Image Quality Rankings

2,400 prompts per cycle across nine active models. Scored on perceptual quality, prompt adherence, and output diversity. No cherry-picked test prompts.

Last run: Aug 5, 2026 2,400 prompts per model 9 models ranked
Top score 88.4
Median 72.1
Range 51.3 - 88.4
Composite: 0.5 quality + 0.3 adherence + 0.2 diversity

Current Rankings

Sorted by composite score. Current cycle as of August 2026. 2,400 structured text prompts covering photographic, illustrative, and abstract categories.

# Model Composite Quality Adherence Diversity Tested
1 nexus-vision-2
88.4
91.2 86.0 84.3 Aug 5, 2026
2 lumina-render-7
85.1
89.4 82.7 79.8 Aug 5, 2026
3 canvas-gen-4
81.6
83.0 84.1 74.2 Aug 5, 2026
4 diffusenet-4
78.3
79.8 77.9 75.4 Aug 5, 2026
5 phasion-3
74.9
76.2 74.3 72.8 Aug 5, 2026
6 inkframe-4
71.0
72.4 73.1 63.7 Aug 5, 2026
7 canvasai-v3
67.5
68.9 66.2 65.4 Jul 29, 2026
8 symbolix-2.5
61.2
62.7 61.4 57.3 Jul 29, 2026
9 pixtral-image-v2
51.3
53.1 50.8 48.4 Jul 29, 2026

How we score image quality

Each model receives 2,400 structured prompts divided across four content categories: photographic scenes, technical diagrams, artistic illustration, and typographic composition. No cherry-picked prompts designed to favor any particular model style.

Composite score weights perceptual quality (50%), prompt adherence (30%), and output diversity (20%). Quality is assessed using a combination of automated no-reference metrics and a held-out panel scoring approach across 200 sampled outputs per model.

Per-category breakdowns and raw evaluation data are available to Pro tier subscribers. Full scoring criteria described on the Benchmarks page.

Access the full image evaluation dataset

Get per-prompt scores, confidence intervals, and category-level breakdowns for all nine models in the current cycle.

See Plans Methodology