We exist to make model quality measurable
Intelligence.AI was founded in 2023 with a single premise: when the organization evaluating a model is the same one selling it, the score is not reliable. We fix that.
What we saw that nobody was saying
In early 2023, Marius Vollberg was running model selection for an AI-native product team. Every vendor benchmark showed their model as clearly superior. When the team ran the same prompts internally, the rankings inverted. Three months later, Priya Natarajan was dealing with the same thing at a different organization: vendor scores that looked authoritative but couldn't survive even basic replication.
The two met at an evaluation methods workshop in San Francisco. By the end of 2023 they had built the first version of what would become the Intelligence.AI benchmark pipeline. Tom Hewson, who had spent years designing evaluation frameworks at an enterprise AI research group, joined in early 2024 as Head of Benchmark Research.
The goal was narrow: publish independent, reproducible scores for models that real teams are actually making procurement decisions about. Not every model. Not every task type. Just the ones that matter most, evaluated rigorously.
The people behind the benchmarks
A small team that cares about measurement quality more than growth metrics.
Marius Vollberg
CEO & Co-Founder
Ran model selection for multiple AI product teams before co-founding Intelligence.AI. Spent two years noticing the same gap between vendor benchmarks and real-world results and decided to fix it rather than tolerate it.
Priya Natarajan
CTO & Co-Founder
Built evaluation infrastructure at a large AI platform before joining as co-founder and CTO. Designed the sandboxed execution environment and the scoring pipeline that powers all three benchmark tracks.
Tom Hewson
Head of Benchmark Research
Spent years designing evaluation frameworks at an enterprise AI research group before joining Intelligence.AI. Leads task set design, contamination detection, and the statistical methodology behind confidence interval reporting.
How we stay operationally independent
Intelligence.AI received angel funding in early 2026 to cover infrastructure and staffing costs. The funding comes with no obligation to produce favorable results for any model provider, and no investor holds board seats that could influence evaluation decisions.
Revenue comes from subscription access to the raw benchmark dataset, the API, and custom evaluation services. Model providers are not customers. This is deliberate: we do not want a financial relationship with the organizations whose products we score.
If you believe a result is wrong, the reproducibility data is available to Pro subscribers. We publish task IDs, sampling seeds, and execution parameters. We have corrected results twice when researchers found genuine errors in our methodology. That transparency is the product.
Questions about methodology?
We answer methodology questions publicly and respond to reproducibility challenges. Reach out via the contact form or read the full evaluation documentation.