Starting today, the Intelligence.AI benchmark data API is publicly accessible on the Free tier. What that means for how teams can integrate scoring into their model selection pipelines.
Since we launched the leaderboards, the most common request we received was not for more models or more task categories. It was for machine-readable access to the data. Teams building internal model selection tooling, evaluation dashboards, or CI pipelines for model regression testing wanted to query our scores programmatically, not copy-paste from a webpage. Today that is no longer necessary.
What the API exposes
The Free tier gives you read access to all published leaderboard scores. You can query current rankings by category (code, image, video), retrieve individual model scores across all three dimensions, and pull the full historical score series for any model we track. All responses return JSON. There is no rate limit on individual requests, only a cap of 100 API calls per day on the Free tier.
The base endpoint structure is intentionally minimal. A request to /v1/leaderboard/{category} returns the ranked list for that category, with each entry containing the model identifier (our internal synthetic name), the composite score, the component subscores, the sample size for that run, the confidence interval at 95 percent, and the evaluation date. A request to /v1/model/{model_id}/history returns the score series for that model across all evaluations we have run, with dates.
This structure reflects a deliberate design choice. We are not exposing raw task-level data in the Free tier, partly for practical reasons (the full task corpus per model is large) and partly because the aggregated scores with confidence intervals are what most integration use cases actually need. The Pro tier adds access to category-level breakdowns and the webhook endpoint for rank-change notifications. The Enterprise tier adds access to private benchmark run results.
Why we opened the Free tier first
We could have launched the API as a paid feature from day one. The leaderboard data has clear value, and we could monetize it immediately. We chose not to, for a specific reason: our credibility depends on the data being verifiable by people who did not pay to see it.
If every team that wanted to cross-check our numbers had to pay first, the implied relationship changes. We are a tiny team, and we cannot be the only people checking whether our evaluation pipeline produces accurate results. Free API access means that evaluation engineers at any organization can pull our data and compare it against their own internal test runs. If our numbers diverge from theirs, we want to hear about it through GitHub issues or direct contact, not discover the discrepancy six months later when a paying customer is frustrated.
There is also a practical benefit: model selection decisions are already being made based on leaderboard data. A team choosing between lumina-7b and cascade-vision for their image pipeline will look at our Image Quality leaderboard. If they can query our API in their selection script and get back structured data rather than parsing HTML, they are more likely to use fresh data rather than a screenshot from two months ago. That freshness matters because we re-evaluate models weekly and scores do change, sometimes significantly after a patch update.
Integration patterns we have seen in early access
We ran a private early-access period with about 40 teams before this public launch. The integration patterns that emerged were not all what we expected.
The most common use case was not the CI regression test we assumed. It was model selection scripting: a team evaluating four or five models for a specific task family would script an API query at the start of their evaluation process to get the current leaderboard state as a prior before running their own internal tests. Our scores served as a first-pass filter, not as a final decision. That is exactly the use case we designed for.
The second most common use case was dashboarding. Several teams integrated our API into internal model monitoring dashboards that tracked which models they were running in production alongside the current leaderboard rank for those models. When a model they were using dropped in rank, the dashboard flagged it for investigation. We added the rank-change webhook to the Pro tier specifically because this pattern kept coming up.
One pattern we did not expect: a small number of teams used the historical score series to track the relationship between model update dates and score changes. They were doing their own analysis of benchmark drift, which maps closely to a piece we published on that topic earlier this year. The history endpoint was not designed with that use case in mind, but it supports it cleanly.
Limits of the current API
We want to be clear about what the API does not do. It does not expose our evaluation code or the task prompts themselves. We publish our methodology in detail on the Benchmarks page, and we describe the task structure, but the specific prompts are a controlled asset. If the task set is public, vendors can optimize against it, and the scores stop meaning what they currently mean. This is the contamination problem in evaluation, and keeping prompts private is one of our main defenses against it.
The API also does not support model submission through it, not yet. If you want to get a model onto the leaderboard, contact us directly. We evaluate models ourselves, which is the whole point: the scores on the leaderboard are produced by our infrastructure, not submitted by the vendor. We are thinking about what a self-serve submission flow would look like for Pro and Enterprise customers, but that work is not done.
Rate limits on the Free tier are real. One hundred calls per day is enough for most dashboard integrations and scripted selection workflows, but if you are building something that needs continuous polling, you will hit the cap. The Pro tier at 10,000 calls per day covers the large majority of automated use cases we have seen.
Getting started
Create a free account at the register page, generate an API key from your dashboard, and pass it as a Bearer token in the Authorization header. The full API reference is linked from your dashboard. We wrote the documentation ourselves and tried to make it short. If something is unclear, the fastest path is [email protected]. We respond quickly because there are only three of us and we all use the product.