Independent rankings for AI models. Measured, not marketed.
Intelligence.AI runs weekly tests across code generation, image quality, and video coherence. Pick a model on actual task scores, not vendor claims.
| # | Model | Score | Pass@1 |
|---|---|---|---|
| 1 | nexus-code-4 | 87.4% | |
| 2 | lumina-7b-turbo | 84.1% | |
| 3 | aurora-pro-2.5 | 80.6% | |
| 4 | orion-70b | 74.8% | |
| 5 | stratum-large-3 | 69.2% |
Three domains, one standard
We measure what matters for teams shipping AI-powered products. Each category uses a distinct evaluation framework built for that output type.
Code Generation
Functional correctness, quality scores, and pass-rate confidence intervals across 400 curated tasks per run. Updated on each model release.
Image Quality
Perceptual quality, prompt adherence, and diversity scoring across 2,400 prompts. Evaluated with both automated metrics and human preference panels.
Video Coherence
Temporal consistency, motion quality, and object permanence scoring across a three-axis framework designed for the unique challenges of video evaluation.
Built for reproducibility
Every evaluation follows a published protocol. You can re-run our benchmarks and arrive at the same numbers.
No vendor affiliation
We do not accept model-provider funding or sponsored rankings. Our only revenue source is subscriptions from teams who use the data, not from the organizations we evaluate.
Published protocols
Each benchmark category has a written methodology document covering task sampling, execution environment, scoring formula, and confidence interval computation. No hidden steps.
Confidence intervals
Single-number scores without uncertainty bounds are not enough. We report 95% confidence intervals on every score so you can compare models where differences are statistically meaningful.
Trusted by teams building with AI
"We evaluated six models for our code review pipeline using Intelligence.AI data. The confidence intervals helped us rule out two that looked equal on vendor dashboards but diverged significantly when task types were broken down."
Ravi S., ML Infrastructure Lead at a B2B software company, from our early-access program
"The image quality leaderboard saved us a three-week evaluation sprint. The prompt adherence axis matched exactly what we needed for our content workflow and the data was reproducible."
Diane K., AI Product Manager at a digital content platform, from our early-access program
Free to explore, affordable to scale
Access full leaderboard data on the Free tier. Pro and Enterprise plans add API access, historical data, and custom evaluations.
Latest from our research
Vendor Benchmark Inflation: Why Self-Reported Scores Keep Missing in Real Tasks
We tracked 18 model releases where vendor-claimed scores diverged from our independent measurements by more than 8 points. Here is what we found.
We Just Opened the Benchmark API to Everyone
Starting today, the Intelligence.AI benchmark data API is publicly accessible on the Free tier.
Inside the Code Evaluation Pipeline: How We Run 400 Tasks Per Model
A walkthrough of the execution infrastructure we built to run code generation benchmarks at scale.
New rankings, when they drop
We re-test weekly. Get the digest in your inbox.
You are on the list. Expect updates when rankings change.