We attempted to reproduce 12 published benchmark results using original methodology descriptions. Eight of twelve required undocumented assumptions to replicate. Here is what we documented.
Reproducibility is a precondition for trust in any empirical claim. A benchmark result that cannot be reproduced by an independent party following the published methodology is not reliable evidence of model capability; it is an anecdote from a specific unspecified setup. We ran a reproduction exercise across 12 published benchmark results, following published methodologies as literally as we could, and tracked what happened. The results were informative but not encouraging.
The exercise and what we were trying to learn
We selected 12 benchmark results from a mix of sources: academic papers that used model evaluation as evidence, technical blog posts from AI research organizations, and one published leaderboard that includes methodology documentation. For each result, we attempted to replicate the reported score by following the described methodology.
Our tolerance for "successful reproduction" was plus or minus 3 points from the reported score. This is a wider tolerance than we would want in principle, but we set it to account for legitimate variation sources: inference temperature variation, minor model version differences within the same named model, and small differences in how we parsed the output. If you cannot reproduce a score within 3 points following the published methodology, the methodology documentation is not sufficient to actually reproduce the result.
We successfully reproduced four of twelve results within the 3-point tolerance using only the published methodology. Four additional results were reproducible after we identified and resolved undocumented assumptions. The remaining four required either significant reverse engineering or were not reproducible within a reasonable tolerance even after investigation.
The categories of undocumented assumptions we found
Across the eight cases that required additional investigation, the undocumented assumptions clustered into five categories.
Prompt format and few-shot construction. Published methodology says "we used 5-shot prompting with examples from the development set." What it does not say: the ordering of the examples, whether examples were selected randomly or matched to the test item in some way, the exact formatting of the separator between examples, and whether the system prompt (if any) was included. These details have non-trivial effects on score. In two of our cases, trying different plausible prompt constructions produced score variation of up to 7 points for the same model.
Model version pinning. When a methodology says "we evaluated model X," it usually does not specify which version of model X in sufficient detail to reproduce the exact binary. For API-accessed models, the published version name may correspond to multiple fine-tuned or patched variants over time. For open-weight models, the version hash is sometimes given, sometimes not. In two cases, the model weight file we could access produced different results than what was published, and we could not determine whether the difference was due to a model version mismatch or something else.
Output parsing and normalization. Code benchmarks score pass/fail on execution output. Language benchmarks often require parsing the model's response to extract the answer before comparison. The details of the parsing logic (how to handle model hedging language, how to extract the numerical answer from a sentence containing it, whether to normalize whitespace and case) can shift scores by several points. Methodology descriptions typically describe the intent of the parsing, not the implementation.
Hardware and numerical environment. This matters more than it should. Floating-point non-determinism across different GPU architectures can affect transformer model outputs even at temperature=0 for some architectures. Two cases showed score differences that we believe are attributable to this effect, though we could not confirm it definitively without access to the original evaluation hardware.
Sampling from the task pool. Several methodologies describe a task pool and state a sample size but do not specify the sampling procedure. "Randomly sampled 500 tasks" does not reproduce unless you know the random seed, the sampling order, or can ensure the samples are stratified equivalently. Two of our eight cases had this issue.
What the four fully irreproducible cases looked like
Two of the four irreproducible cases involved models that are no longer accessible in the same form as when the benchmark was run. One was a model whose API access has since changed. The other was an open-weight model where the published weight hash does not match any available checkpoint we could find.
One case involved a score gap that we could not close even after reconstructing what we believe were the correct settings. Our best attempt was 9 points below the published result. We do not attribute this to dishonesty; we suspect there is an undocumented configuration detail we are missing. But we could not find it.
The fourth case was a benchmark result published without any methodology documentation beyond the model name and task name. We tried four plausible evaluation setups and produced scores ranging from 61 to 78 for the same model on what should have been the same task. The published result was 84. Without methodology documentation, we have no way to understand this gap.
What good reproducibility documentation looks like
The bar we set for our own benchmark documentation is: another team with API access to the model should be able to reproduce our score within 2 points by following our published methodology alone. This requires documenting: the exact model identifier and version hash where available, the inference configuration parameters (temperature, max tokens, top-p, system prompt verbatim), the prompt format including exact separators and few-shot example selection logic, the output parsing procedure with code-level precision (we link the parser), the random seed and sampling procedure for any random selection, and the execution environment including language runtime versions.
We make our scorer code available on request. We log evaluation runs so we can re-examine them. We accept methodology disputes from researchers who attempt to reproduce our results and find unexplained discrepancies. This is not a claim that our methodology is perfect; it is a claim that it is auditable.
Why this matters beyond benchmarks
Reproducibility in model evaluation is a specific instance of a broader problem in applied machine learning: the gap between what a result claims and what another team can verify. When evaluation results are not reproducible, the model selection decisions based on them have unknown reliability. A team that chooses nexus-code-4 over an alternative based on a benchmark score that cannot be reproduced has made a decision based on evidence with unknown validity.
The benchmark publishing ecosystem has not developed the same reproducibility norms as, say, clinical trials or semiconductor performance specifications, despite the fact that the downstream decisions based on model benchmarks are increasingly consequential. We do not think that gap will close quickly, but we think operating a benchmark platform as if reproducibility standards already applied is the right approach. The 8-of-12 finding from this exercise is the reason our own methodology documentation is as detailed as it is.