Benchmark Lab Methodology

Designing Benchmarks for Reasoning Models: What Standard Task Sets Miss

Tom Hewson  · 

Abstract branching decision tree representing the complexity of benchmark design for reasoning models

Chain-of-thought models expose weaknesses in benchmarks designed before extended reasoning became mainstream. We rewrote our task structure for the new generation.

Standard benchmark task sets were designed to probe model output directly. You give the model a prompt, it returns a completion, you score the completion against a reference answer. That evaluation paradigm works reasonably well for models that operate in a single generation pass. It produces increasingly misleading results when applied to models that emit extended reasoning traces before arriving at a final answer.

We spent three months rebuilding our task structure after running into this problem on our own leaderboard. This is a description of what we found and how we responded.

The scoring problem with chain-of-thought outputs

The most immediate symptom is score ambiguity. A reasoning model might produce 800 tokens of working-through-the-problem before writing the final answer. If your scorer is looking for the reference answer anywhere in the completion, it will find it in the chain-of-thought regardless of whether the model eventually reaches the correct conclusion. A model that reasons its way to the wrong answer, notices the error, and self-corrects before the final response will score the same as a model that goes straight to the right answer. But those are meaningfully different behaviors.

We saw this concretely with orion-r1, a reasoning model we added to our code leaderboard in early testing. Under our original pass-rate scorer, which checked whether the correct output appeared in the completion, orion-r1 scored 79.2 on our tier-2 algorithm tasks. When we tightened the scorer to evaluate only the final answer block and ignore the reasoning trace, the score dropped to 68.4. The 10.8-point gap was not noise. The model was frequently reasoning to a correct intermediate step and then proceeding to a different final answer. Our original scorer was rewarding the intermediate correct step, not the actual output.

The fix for this specific problem was straightforward: define a structured output format for task responses where models must place their final answer in a delimited block, and score only that block. For code tasks, this means the final code solution must appear in a code fence tagged with the language, and only that block is executed against the test suite. For classification and extraction tasks, the final answer must appear in a specific field. Models that do not comply with the output format structure are scored on the raw completion, which tends to produce lower scores because it includes the reasoning noise.

Multi-step tasks and intermediate state

The deeper structural problem is that standard benchmark tasks are single-step: one input, one correct output. Reasoning models are often most useful on problems that require intermediate state, planning, and sequential decision-making. A benchmark that only tests one-shot input-output pairs will systematically undervalue reasoning models relative to their actual utility on hard problems, and will fail to distinguish good reasoning from lucky single-step guessing.

We added a category of multi-step tasks to our code evaluation pool specifically to address this. A multi-step task has a defined problem state, a sequence of required intermediate operations, and a final output that cannot be reached by luck on a single generation pass. An example: given a partially specified data structure with three invariants, implement the insert and delete operations such that all invariants hold after each modification, and provide test output for five specific operation sequences. This cannot be solved by memorizing a pattern. A model must plan the operations, implement them with the invariants in mind, and trace the state through the required sequences.

Multi-step tasks make up 30 percent of our tier-3 task pool. They are more expensive to score (we have to trace the state through each step rather than just checking the final output) and they require more careful prompt construction to specify the required intermediate outputs. But they are also the tasks that best differentiate reasoning-capable models from models that have strong pattern-matching but limited planning capability.

Token budget effects

An underappreciated complication with reasoning model evaluation is token budget sensitivity. Some reasoning models perform meaningfully better when given a larger token budget for their chain-of-thought, and meaningfully worse when that budget is constrained. If you evaluate two models under different token budget conditions, you are not comparing the models. You are comparing the models plus their allowed thinking space.

We now enforce a fixed token budget for all code evaluation runs: 4,000 tokens for the reasoning trace and final response combined, applied uniformly across all models. This is a constraint that is tighter than what most real-world deployments use, and it means our scores represent a budget-constrained capability rather than the maximum possible performance of a well-configured reasoning model. We are explicit about this in the leaderboard methodology note.

We made this choice because the alternative, allowing each model to use its preferred token budget, would produce incomparable scores. A reasoning model using 8,000 tokens of thinking space is not directly comparable to a standard model using 200 tokens of response space. The fixed budget makes the scores comparable. It trades some absolute ceiling measurement for comparability, which is the right tradeoff for a leaderboard.

What standard task sets miss about failure modes

There is a category of problem that comes up frequently in practice where standard benchmarks provide almost no coverage: tasks where the model must recognize that a specification is underspecified or contradictory and respond accordingly rather than generating a confidently wrong answer.

Reasoning models are often particularly bad at this. A model with strong planning capability will sometimes construct an elaborate solution to an underspecified problem that is internally consistent but answers the wrong question because the model never flagged the ambiguity. Standard benchmarks do not include ambiguous or contradictory specifications because they are hard to score. Our reference-answer-based scoring cannot easily reward a model that says "this specification has two possible interpretations, and here is what I would do for each" over one that silently picks an interpretation and returns a plausible-looking answer.

We added a small category of ambiguity-detection tasks to our evaluation pool, where the correct response type is a clarifying question or an explicit statement of the assumed interpretation before proceeding. These are binary scored: either the model flags the ambiguity, or it does not. Currently 20 tasks out of 400 are in this category, and the distribution of correct responses across models is quite different from the distribution on normal tasks. Some models that rank highly on standard functional tasks rank much lower on ambiguity detection. That divergence is information.

Limitations of the new task structure

We are not claiming that our revised task structure fully characterizes reasoning model capability. The categories we test (algorithm implementation, multi-step code construction, ambiguity detection) cover a portion of the use cases where reasoning models provide real value. Long-horizon planning tasks, tasks requiring external tool use, and tasks where the quality of the reasoning trace itself matters (not just the final answer) are not well-represented in our current pool.

We also have not resolved the question of how to score partial credit on multi-step tasks where a model gets the planning right but makes an implementation error in the last step. We currently score these as full failures, which is probably too harsh. A future version of the scoring will try to assign credit to correct intermediate steps. The challenge is that partial credit scoring introduces judgment calls that are hard to make consistent at scale.

Read the benchmark methodology

Our full task design documentation, scoring criteria, and reproducibility standards.

View Methodology