Benchmark Lab Methodology

Inside the Code Evaluation Pipeline: How We Run 400 Tasks Per Model

Priya Natarajan  · 

Abstract pipeline diagram showing the stages of automated code evaluation

A walkthrough of the execution infrastructure we built to run code generation benchmarks at scale: task sampling, sandboxed execution, and pass-rate aggregation.

Running 400 code generation tasks against a model sounds straightforward until you try to do it reproducibly, at weekly cadence, across 28 models at the same time. This post describes the pipeline we built to make that work. It is not a research contribution. It is an engineering diary of the decisions we made and the mistakes we fixed along the way.

Task composition and sampling

Our code evaluation task pool currently holds 2,400 items across seven task categories: function completion, algorithm implementation, bug fixing, code translation between languages, documentation generation, unit test writing, and refactoring. We do not use the full pool for every evaluation run. For the weekly leaderboard update, we sample 400 tasks using a stratified procedure that maintains fixed proportions across task categories and difficulty tiers.

Difficulty tiers are assigned during task construction, not by model performance after the fact. We classify each task at intake based on three factors: the number of independent logical conditions the solution must satisfy, the size of the solution space (rough count of plausible correct approaches), and the presence of edge cases that require explicit handling beyond the obvious path. A tier-1 task is something like "write a function that returns the sum of a list." A tier-3 task is something like "implement a cache-aware merge sort for a custom linked list type that maintains sorted order under concurrent inserts." The tier distribution in our 400-task samples is 40 percent tier-1, 40 percent tier-2, 20 percent tier-3. We maintain this distribution because a flat tier distribution would compress score differences at the top of the leaderboard into noise.

Task sampling uses a fixed random seed per evaluation cycle. The seed is derived from the ISO week number and the year, which means every model in a given week's run gets evaluated on exactly the same 400 tasks. This is the foundation of apples-to-apples comparison. We log the seed and the exact task IDs included in each run, so any score can be traced back to the specific task set that produced it.

The execution environment

Code execution is the part of the pipeline that causes the most operational complexity. You cannot run arbitrary model-generated code on shared infrastructure without either sandboxing it properly or accepting the risk that a model will output something that modifies the filesystem, makes network calls, or runs an infinite loop that blocks the queue.

We use container-based isolation for each task execution. Each code sample gets its own container with a read-only filesystem, no network access, resource limits (256 MB memory, 2 CPU seconds of wall time, 30 MB process memory), and a forced termination at the timeout boundary. The timeout is the most frequently triggered limit. Long-running loops and quadratic algorithms in tier-2 tasks hit it regularly. A timeout is scored as a failure for that task, the same as incorrect output.

Container startup overhead was a real engineering problem at launch. Cold-starting a container per task added about 1.2 seconds per execution, which across 400 tasks per model and 28 models per weekly run amounted to significant total wall time. We moved to a pool of warm containers with reset between uses rather than full teardown and rebuild. Resetting means forcibly terminating all processes, restoring the filesystem to its baseline state from a snapshot, and clearing the memory space. The reset operation takes about 180 milliseconds. At scale, this is roughly a 7x reduction in per-task overhead compared to cold-start containers.

Language support is currently Python, JavaScript (Node.js), and Java. We made a deliberate choice to constrain the language set rather than try to cover everything. Multi-language support requires maintaining separate runtime environments, separate test runners, and separate correctness heuristics. Adding a new language to the pipeline is a week of work, minimum, to do it right. We will expand, but we are not going to add a language until we can test it properly.

Automated scoring and pass-rate aggregation

Pass-rate measurement has two components: functional correctness and output format compliance. Functional correctness is checked against a test suite per task. Each task in our pool has between 5 and 20 test cases, including public cases and private hidden cases. A solution must pass all test cases to score a pass on that task. Partial credit would introduce subjectivity in test case weighting, so we do not use it. Either the function returns the correct output for every input in the test suite, or it does not pass.

Format compliance is a secondary check. If a task asks for a function with a specific signature and the model returns a class, a script, or prose explaining how it would write the function, that response fails format compliance and is scored as a non-pass regardless of whether the underlying logic is correct. This sounds harsh, but it reflects real-world conditions. A developer querying a model API for a function and receiving back an essay cannot use that response directly.

The aggregate pass rate for a model on a given run is the number of task passes divided by 400. We report this as the raw score. We also compute bootstrapped confidence intervals at 95 percent using 1,000 resamples of the 400-task result set. This tells you how much of the score gap between two adjacent models is attributable to statistical noise in task sampling versus a real performance difference. Two models separated by less than 2 points on a 400-task sample are often within each other's confidence intervals, and we flag that explicitly in the leaderboard display.

Weekly refresh and score versioning

We re-run every model each week. This is more expensive than updating only when models release new versions, but it is necessary for two reasons. First, model providers sometimes push silent patch updates that change model behavior without announcing a version change. We documented four cases of this in our benchmark drift article. Second, our task pool rotates a small percentage of tasks each cycle to prevent over-optimization, and the score on a stable model should track that rotation so we are measuring a clean comparison.

Each evaluation run produces a versioned score record: model identifier, evaluation date, task set seed, raw score, confidence interval bounds, and component subscores by task category and difficulty tier. We keep all historical records. When you query the API for a model's score history, you are getting the full sequence of versioned records, not just the current value.

What the pipeline does not do

This infrastructure handles functional correctness. It does not handle code quality: readability, maintainability, idiomatic use of language features, or robustness to edge cases not covered by the test suite. A model can pass all 400 tasks by generating correct but unreadable one-liners or by using brute-force approaches that work on the given test suite but would fail on slightly different inputs. We are building a complementary quality scoring dimension for exactly this reason. The functional pass rate is necessary information; it is not sufficient information for a real model selection decision. That limitation is intentional to be explicit about, not something we want to paper over.

Explore the code leaderboard

See current rankings from the pipeline described in this article.

View Code Leaderboard