We re-ran seven models after their providers released silent patch updates. In four cases, scores shifted by more than three points without any public announcement.
One of the things our weekly re-evaluation cadence is designed to catch is model behavior changes that happen between named version releases. Providers update model weights, system prompts, safety filters, and inference configuration without necessarily releasing a new named model or publishing a changelog. These silent updates can change benchmark scores significantly. Over the first five months of our leaderboard, we documented seven cases where a model's behavior changed in ways detectable by our benchmark, and four of those cases involved score shifts of more than three points.
How we detect drift
The weekly re-evaluation produces a score record per model per cycle. We compute the change between consecutive weekly scores and flag any shift greater than 1.5 points (our threshold for "likely more than sampling noise" based on our confidence interval analysis) as a drift event. A drift event is investigated before the updated score is published, to distinguish genuine model behavior change from evaluation artifact.
An evaluation artifact is when our pipeline produces a different score for the same underlying model behavior, due to changes in our task sampling, execution environment, or scoring logic. We run a set of stability tests designed to detect this: we re-run a fixed held-out set of 50 tasks on the same model version before and after any suspected drift, using a frozen task set that we never rotate. If the held-out set score is stable while the sampled-set score changed, the drift is an artifact of our sampling and we do not publish the change as model drift. If the held-out set score also changed, the model behavior itself changed.
All seven cases we document here passed the held-out test. The behavior change was in the model, not our pipeline.
The four significant cases
We are not naming providers for these cases because our relationship with providers is built on the premise that we evaluate independently without their involvement, and attribution of specific patch failures creates a dynamic that works against honest evaluation. We describe the patterns instead.
Case 1: A code generation model dropped 5.2 points on our tier-2 algorithm tasks following a patch update. The drop was concentrated in a specific subcategory: recursive tree traversal problems. The model had previously been reliable on these tasks, and post-patch it began producing iterative solutions to recursive problems, which passed fewer of our test cases (our test suite included recursive-output verification for several tree tasks). We do not know whether the patch was targeting recursion behavior or whether the change was a side effect of something else.
Case 2: An image generation model improved by 4.8 points on our prompt adherence score following a patch. This was a positive drift event. The improvement was consistent across multiple prompt categories, suggesting a deliberate tuning effort. The provider published no announcement, but the consistent directional improvement across categories rules out chance. The model moved from rank 6 to rank 4 in our image leaderboard following this update.
Case 3: A video model changed its output format in a way that broke our automated parser. The model started wrapping its video generation parameters in a different JSON structure than it previously used. This initially appeared in our monitoring as a large negative drift event. When we investigated, we found the parser failure and corrected it. After correction, the score was actually 2.1 points higher than the pre-patch score. This case is a reminder that format changes are the hardest to detect reliably, since the failure mode can look identical to a genuine score drop until you inspect the raw outputs.
Case 4: A code generation model dropped 3.4 points on function completion tasks after a patch. In this case, the model appeared to have become more verbose: it started including explanatory comments in the generated code, which changed the timing profile of execution and caused more solutions to hit our 2-second wall-clock timeout. The underlying algorithms were often correct; they just ran slower because of the additional comment processing overhead and more defensive code patterns. Whether "adds explanatory comments" is a quality improvement or degradation depends entirely on your use case, but from a functional pass-rate perspective it registered as a loss.
The transparency problem
The core issue these cases illustrate is that model versions and model behaviors are not synchronized in the current ecosystem. A model accessed via API may change its behavior without any versioning signal that would allow a downstream user to detect the change. For a team that built a workflow around a specific model's behavior, a silent update is an invisible change to a dependency they cannot pin or roll back.
This is not a novel problem; it exists in other software contexts too. But in ML models, the difficulty of detecting behavioral change is higher than in traditional software. A code library that changes its output format will immediately break integrations in observable ways. A language model that becomes slightly more verbose or slightly more conservative on a specific task type may not break anything obviously while still producing meaningfully different outputs that affect downstream quality.
Our weekly re-evaluation is a partial mitigation. If you subscribe to rank-change notifications through our API, you will know within a week when any model in our leaderboard shows a significant score shift, which is a proxy signal for behavioral change. It is not a complete solution, but it is currently better than having no signal at all.
What we are building to address this
We are building a behavioral fingerprinting feature that will track a fixed set of model outputs over time and alert when those outputs change. The idea is to run a small set of deterministic prompts (prompts designed to produce the same output reliably if the model is unchanged) against each model on every evaluation cycle. Changes in those fixed outputs are strong evidence of model behavior change independent of our sampled benchmark tasks.
The challenge is that truly deterministic prompts are harder to design than they sound. Many prompts that seem deterministic produce variation due to temperature-based sampling. We use temperature=0 for this probe set, which reduces but does not eliminate stochasticity for all model architectures. The fingerprinting feature is in testing and is not yet in the production pipeline. When it ships, it will show up as an additional data field in the API response for each model: a behavioral stability indicator based on the output stability of our fixed probe set.
A note on interpretation
Benchmark drift is not inherently bad. Case 2 above was a genuine model improvement. We are not arguing that providers should avoid updating their models or that all drift is a problem. We are arguing that users who rely on model behavior in production deserve a signal when that behavior changes, and that silent updates without any public notice create real operational problems for teams that have built workflows depending on specific model characteristics.