Benchmark Lab Methodology

Code Correctness vs. Code Quality: Why Pass Rate Alone Is Not Enough

Priya Natarajan  · 

Abstract split comparison between a passing test suite and quality metrics on generated code

A model can pass functional tests while generating brittle, unreadable code. We are building a complementary quality dimension into our code benchmark and here is how.

Pass rate is the dominant metric in code generation benchmarks because it is easy to measure and hard to game without actually solving the problem. Either the code produces the correct output or it does not. This clean binary property makes pass rate the right foundation for any code benchmark.

It is not sufficient on its own. This post explains the problem, describes the quality signals we are adding to our code evaluation, and is honest about the parts we have not solved yet.

The problem with pass-rate-only evaluation

Consider a function that correctly computes the longest common subsequence of two strings. There are several ways to implement this correctly. A clear, maintainable implementation using dynamic programming with explicit variable names takes about 20 lines. A one-liner using a recursive lambda with list comprehensions might also pass all test cases while being essentially unreadable to anyone who did not write it. A third implementation might use a brute-force exponential-time approach that passes all test cases in the benchmark because the test inputs are small, while being completely unusable for any real input of meaningful size.

All three pass the correctness check. A model that consistently produces the third type is not useful for production code generation, even if its benchmark pass rate is high. Pass rate measures whether the model can produce code that gets the answer right under test conditions. It does not measure whether the code is something a human engineer could review, debug, or maintain.

This distinction matters more now than it did two years ago. Code generation models are being used not just to write throwaway scripts but to write production code that will be maintained by teams over time. A model that produces correct but unreadable or architecturally brittle code creates technical debt that is often not visible at the time of generation.

Four quality dimensions we are measuring

We identified four quality dimensions where objective measurement is feasible without requiring human review for every generated solution. These dimensions are complementary to correctness, not replacements for it. We only compute quality scores for solutions that pass correctness first.

Cyclomatic complexity relative to task difficulty. Cyclomatic complexity (roughly, the number of independent paths through a piece of code) is a proxy for maintainability. A function with complexity 3 that solves a tier-1 task is expected. A function with complexity 15 that solves a tier-1 task is doing something unnecessarily complicated. We compute the ratio of measured complexity to expected complexity for the task difficulty tier, and penalize outliers on both the high (overcomplicated) and low (suspiciously clever) ends.

Algorithmic efficiency class. For tasks where we know the theoretically optimal time complexity, we run solutions against input sizes that distinguish complexity classes: O(n) versus O(n log n) versus O(n^2) versus O(2^n). We test each passing solution with inputs of size 10, 100, 1,000, and 10,000, measure execution time, and fit the scaling curve. A solution that claims to solve an O(n log n) task in linear time is suspicious; one that solves it in quadratic time is technically correct but will fail in production. This check takes more compute per solution but is fully automatable.

Security anti-pattern detection. We run static analysis on generated Python code to detect common security anti-patterns: SQL string formatting instead of parameterized queries, use of eval() or exec() on variable input, subprocess calls with shell=True, insecure deserialization patterns. A solution that passes functional tests while containing an obvious SQL injection vulnerability is not a good output even if it is technically correct for the given test cases. We flag solutions with detected anti-patterns and score them lower on the quality dimension.

Variable naming and documentation density. This is the softest of our four measures. We count the proportion of single-letter variable names, check that function signatures include type annotations for Python functions, and check whether the function has a docstring. None of these are absolute quality signals; there are contexts where single-letter variable names are idiomatic (loop indices, coordinate pairs in mathematical code). We treat this as a style signal rather than a correctness-adjacent signal and weight it lightly in the composite quality score.

The composite quality score

Our current weighting for the quality composite is: complexity ratio 30 percent, efficiency class 40 percent, anti-pattern score 20 percent, naming and documentation 10 percent. We weighted efficiency most heavily because it is the quality failure that most reliably causes real problems in production (a too-clever variable name is annoying; a quadratic algorithm on large inputs is a service incident). The weights were calibrated against a set of manually reviewed solutions from our internal benchmark runs, asking "which of these would a senior engineer prefer in a code review?" The calibration set was 200 solutions, and we validated the composite score against holdout manual reviews.

The composite quality score is a separate metric from pass rate. We do not combine them into a single number for the leaderboard. We report pass rate and quality score as two independent dimensions, and we let users weight them according to their use case. A team building a script-automation tool might weight pass rate heavily. A team building a code generation feature that will insert code into production codebases should care more about quality.

What we have not solved

Two quality signals that matter in practice are not yet in our pipeline. The first is idiomatic language use: code that is technically correct and reasonably efficient but written in a style that is inconsistent with the conventions of the language. Python code written like C++, for example, will pass correctness and efficiency checks while being awkward for Python developers to maintain. Measuring idiomaticity automatically is hard; the best approaches involve comparing generated code to a reference corpus of high-quality idiomatic code, and building that reference corpus for multiple languages is non-trivial work.

The second is architecture-level quality for multi-function tasks. Our efficiency and complexity measures work well for single functions. For tasks that require writing multiple related functions (a small module, a class with methods), there are quality properties at the architectural level that our per-function measures miss: cohesion, appropriate separation of concerns, consistent interface design. Measuring these requires more structured analysis than our current pipeline supports. This is on the roadmap but is not solved.

Current status and access

The quality dimension is currently in evaluation on a portion of our code task set. It is not yet reflected in the main leaderboard scores, which still show pass rate as the primary metric. We are running both metrics in parallel and will switch to the two-dimension display when we have enough data to validate the quality scores have adequate test-retest reliability across weekly runs. We expect that to be ready within the next evaluation cycle. Pro and Enterprise API users will see the quality sub-scores in the API response before they appear in the public leaderboard display.

Code leaderboard with quality dimensions

Our code generation rankings include correctness, readability, and security sub-scores alongside the overall rating.

View Code Leaderboard