Code Generation Rankings
400 tasks per model. Sandboxed execution. Scored on correctness, code quality, and edge-case handling. Updated weekly from live test runs.
Current Rankings
Sorted by composite score. Click column headers to re-sort. Scores rounded to one decimal.
| # | Model | Composite | Pass@1 | Quality | Edge | Tested |
|---|---|---|---|---|---|---|
| 1 | nexus-code-4 | 87.4% | 94.1 | 88.3 | Aug 7, 2026 | |
| 2 | lumina-7b-turbo | 84.1% | 91.2 | 85.9 | Aug 7, 2026 | |
| 3 | aurora-pro-2.5 | 81.7% | 88.0 | 82.4 | Aug 7, 2026 | |
| 4 | lumina-7b-mini | 78.3% | 84.9 | 79.8 | Aug 7, 2026 | |
| 5 | nexus-code-4s | 76.9% | 82.3 | 77.2 | Aug 7, 2026 | |
| 6 | aurora-flash-2 | 73.5% | 80.6 | 74.1 | Aug 7, 2026 | |
| 7 | orion-70b | 71.4% | 77.8 | 71.9 | Aug 7, 2026 | |
| 8 | helix-3 | 69.2% | 75.4 | 70.3 | Aug 7, 2026 | |
| 9 | stratum-large-3 | 66.8% | 73.1 | 67.0 | Jul 31, 2026 | |
| 10 | codestream-32b | 63.9% | 69.8 | 64.2 | Jul 31, 2026 | |
| 11 | korex-coder-v3 | 58.7% | 65.0 | 59.8 | Jul 31, 2026 | |
| 12 | orion-8b | 48.3% | 54.7 | 49.9 | Jul 31, 2026 |
How we score code generation
Each model runs the same 400 tasks across five categories: algorithmic problems, data transformation, API integration stubs, debugging challenges, and documentation generation. Tasks are drawn from a private set refreshed quarterly to reduce contamination risk.
The composite score weights pass@1 (60%), code quality assessment (25%), and edge-case handling (15%). Quality is measured by a separate static analysis pass scoring readability, correctness of types, and absence of common anti-patterns.
Confidence intervals at 95% are available via the Pro tier API. Full methodology details are on the Benchmarks page.
Access the full benchmark dataset
Download raw scores, confidence intervals, and per-task breakdowns. Pro and Enterprise plans include API access for pipeline integration.