Image Quality Rankings
2,400 prompts per cycle across nine active models. Scored on perceptual quality, prompt adherence, and output diversity. No cherry-picked test prompts.
Current Rankings
Sorted by composite score. Current cycle as of August 2026. 2,400 structured text prompts covering photographic, illustrative, and abstract categories.
| # | Model | Composite | Quality | Adherence | Diversity | Tested |
|---|---|---|---|---|---|---|
| 1 | nexus-vision-2 | 91.2 | 86.0 | 84.3 | Aug 5, 2026 | |
| 2 | lumina-render-7 | 89.4 | 82.7 | 79.8 | Aug 5, 2026 | |
| 3 | canvas-gen-4 | 83.0 | 84.1 | 74.2 | Aug 5, 2026 | |
| 4 | diffusenet-4 | 79.8 | 77.9 | 75.4 | Aug 5, 2026 | |
| 5 | phasion-3 | 76.2 | 74.3 | 72.8 | Aug 5, 2026 | |
| 6 | inkframe-4 | 72.4 | 73.1 | 63.7 | Aug 5, 2026 | |
| 7 | canvasai-v3 | 68.9 | 66.2 | 65.4 | Jul 29, 2026 | |
| 8 | symbolix-2.5 | 62.7 | 61.4 | 57.3 | Jul 29, 2026 | |
| 9 | pixtral-image-v2 | 53.1 | 50.8 | 48.4 | Jul 29, 2026 |
How we score image quality
Each model receives 2,400 structured prompts divided across four content categories: photographic scenes, technical diagrams, artistic illustration, and typographic composition. No cherry-picked prompts designed to favor any particular model style.
Composite score weights perceptual quality (50%), prompt adherence (30%), and output diversity (20%). Quality is assessed using a combination of automated no-reference metrics and a held-out panel scoring approach across 200 sampled outputs per model.
Per-category breakdowns and raw evaluation data are available to Pro tier subscribers. Full scoring criteria described on the Benchmarks page.
Access the full image evaluation dataset
Get per-prompt scores, confidence intervals, and category-level breakdowns for all nine models in the current cycle.