Leaderboard

Code Generation Rankings

400 tasks per model. Sandboxed execution. Scored on correctness, code quality, and edge-case handling. Updated weekly from live test runs.

Last run: Aug 7, 2026 400 tasks per model 12 models ranked
Top score 91.2
Median 74.6
Range 52.1 - 91.2
Composite score: 0.6 pass@1 + 0.25 quality + 0.15 edge

Current Rankings

Sorted by composite score. Click column headers to re-sort. Scores rounded to one decimal.

# Model Composite Pass@1 Quality Edge Tested
1 nexus-code-4
91.2
87.4% 94.1 88.3 Aug 7, 2026
2 lumina-7b-turbo
88.7
84.1% 91.2 85.9 Aug 7, 2026
3 aurora-pro-2.5
85.3
81.7% 88.0 82.4 Aug 7, 2026
4 lumina-7b-mini
82.1
78.3% 84.9 79.8 Aug 7, 2026
5 nexus-code-4s
80.4
76.9% 82.3 77.2 Aug 7, 2026
6 aurora-flash-2
77.8
73.5% 80.6 74.1 Aug 7, 2026
7 orion-70b
75.2
71.4% 77.8 71.9 Aug 7, 2026
8 helix-3
73.0
69.2% 75.4 70.3 Aug 7, 2026
9 stratum-large-3
70.5
66.8% 73.1 67.0 Jul 31, 2026
10 codestream-32b
67.3
63.9% 69.8 64.2 Jul 31, 2026
11 korex-coder-v3
62.4
58.7% 65.0 59.8 Jul 31, 2026
12 orion-8b
52.1
48.3% 54.7 49.9 Jul 31, 2026

How we score code generation

Each model runs the same 400 tasks across five categories: algorithmic problems, data transformation, API integration stubs, debugging challenges, and documentation generation. Tasks are drawn from a private set refreshed quarterly to reduce contamination risk.

The composite score weights pass@1 (60%), code quality assessment (25%), and edge-case handling (15%). Quality is measured by a separate static analysis pass scoring readability, correctness of types, and absence of common anti-patterns.

Confidence intervals at 95% are available via the Pro tier API. Full methodology details are on the Benchmarks page.

Access the full benchmark dataset

Download raw scores, confidence intervals, and per-task breakdowns. Pro and Enterprise plans include API access for pipeline integration.

See Plans Methodology