When the organization measuring model performance is the same one selling the model, there is a conflict. We wrote this to explain why we think third-party evaluation matters.
We have a stake in making the case for independent benchmarks because we built one. That makes this argument suspect, and we want to acknowledge that directly before making it. But the conflict does not make the argument wrong, and the argument does not depend on our own platform being good. It depends only on the claim that evaluation and commercialization are genuinely in tension when combined in the same organization.
The structural conflict
A model provider that publishes its own benchmark results faces a set of decisions that are all downstream of the same conflict of interest. Which benchmark tasks to include. What inference configuration to use. When to run the evaluation and whether to re-run it if results are disappointing. How to present the numbers and which comparisons to highlight. At every decision point, the commercially rational choice and the epistemically honest choice diverge. This does not mean vendor benchmarks are dishonest. It means they are made by people with incentives that do not align with maximizing the accuracy of the reported score.
This structural problem is not unique to AI. It appears wherever the producer of a thing also controls the measurement of that thing. We have learned, across many industries and over many decades, that self-certification in high-stakes domains produces worse outcomes than independent certification. The stakes in model selection for production AI systems are rising fast, and the current norm of vendor-reported benchmarks as the primary evidence for model quality is not going to hold.
What "independent" actually requires
Independence has specific practical requirements, not just an organizational structure. An entity can be legally separate from model providers while still having evaluation relationships that compromise its independence. The critical independence requirements are: the evaluating organization cannot take revenue from the providers whose models it evaluates; it cannot accept early access to models in exchange for favorable treatment or embargo timing; it cannot allow providers to pre-review scores before publication; and it cannot design its evaluation methodology in consultation with providers.
We observe all four requirements. Model providers are not our customers. We do not have revenue relationships with any provider. We do not accept early access for our models. Scores are published on our schedule, not subject to provider review periods. Our methodology is designed and revised by our team without provider input. This is the minimum baseline for independence, not a gold standard. We are a small team with limited resources, and there are aspects of evaluation rigor that well-resourced academic groups can achieve that we currently cannot. But the absence of a revenue relationship with the evaluated entities is the non-negotiable foundation.
Why the "test it yourself" response is insufficient
The standard pushback against the case for independent benchmarks is that teams should evaluate models on their own specific tasks rather than relying on any benchmark. This is correct advice for final selection decisions on high-stakes use cases. It is not a sufficient response to the case for independent benchmarks, for two reasons.
First, most teams do not have the resources to run comprehensive evaluations across the full field of available models before narrowing their selection. A useful benchmark is a filter: it reduces the field to a manageable set of candidates for deeper evaluation. A vendor-reported benchmark is a poor filter because you do not know how much it has been optimized for favorable presentation. An independent benchmark using fixed, published methodology is a better filter because the optimization incentive is absent.
Second, specific task evaluation does not address the information asymmetry that affects the market as a whole. When evaluation methodology is not standardized or independent, the signal quality of capability claims degrades across the ecosystem. A researcher or engineer who wants to understand the current state of the field relies on comparative information. If all available comparative information is vendor-produced, their picture of the field is systematically distorted in ways that are hard to detect from any single evaluation they run themselves.
The academic benchmark alternative
Academic research groups have produced some of the best benchmark infrastructure in the field, and some of those benchmarks do achieve genuine independence from commercial providers. But academic benchmarks face their own structural problems that limit their usefulness for production model selection.
Academic benchmarks tend to be static. They are designed, published, and then held constant to allow longitudinal comparison. That is the right design for research: you want to know whether progress is real, and you can only judge that against a fixed reference. But a static benchmark degrades over time as training data contamination accumulates and as models are specifically optimized for it. The benchmark stops measuring general capability and starts measuring benchmark-specific performance. The history of any prominent benchmark shows this pattern clearly.
Academic groups also tend not to run the models they evaluate at the cadence that production users need. A benchmark paper that is published once and includes model evaluations from six months ago is not useful for selecting among models available today. Weekly re-evaluation at the cadence of model releases is an operational commitment that academic benchmark projects rarely maintain.
What we are not claiming
We are not claiming that vendor benchmarks should be disregarded or that they contain no information. Vendors run evaluations with access to their models' capabilities that no external party can fully replicate. A vendor knows the range of configurations their model was designed to operate under, the task types it was specifically trained on, and the inference settings that best reflect intended use. That knowledge, if honestly applied, produces benchmark results that are informative about specific capabilities under specific conditions.
The problem is not that vendor benchmarks are uninformative. The problem is that they are optimized for favorable presentation in ways that the reader cannot assess, and they are not comparable across vendors in ways that the reader can trust. Independent benchmarks under fixed, published methodology fill a specific gap: comparability under controlled conditions. They are not a replacement for all other forms of evaluation. They are a specific tool for a specific problem.
The long run
The case for independent benchmarks is strongest as a long-run argument. In any individual evaluation cycle, a motivated vendor can produce results that beat an independent evaluation simply by optimizing harder for the task distribution. The independent evaluator cannot permanently stay ahead of that optimization. What independent evaluation provides, consistently over time, is a ground truth that does not drift with commercial incentives. The value compounds as the benchmark establishes a track record and as users learn to trust the comparison basis.
We are at the early stage of that compounding. The benchmark we have built is not the benchmark we want to have in two years. The methodology will improve, the task coverage will grow, and the number of models we track will expand. The thing that will not change is the structural separation between our revenue sources and the models we evaluate. That separation is the product we are actually selling, even when the immediate deliverable looks like a ranked list of model scores.