FrontierSWEV2

Benchmarking software engineering skill at the edge of human ability.

By

Leaderboard

Scores across all 34 tasks. Each model runs 5 trials per task with a 20-hour budget.

#ModelScore
1
GPT-6 Astra
proximus
65.5%
±8.9
2
Claude Opus 5.5
proximus
62.3%
±9.8
3
GLM-5.3
proximus
30.2%
±11.5
4
Grok 4.7
proximus
29.5%
±12.6
5
Kimi K3
proximus
25.9%
±11.8
6
Gemini 3.7 Flash
proximus
20.3%
±10.5
7
Qwen3.8-Max
proximus
15.8%
±7.8
8
DeepSeek V4 Flash Vision Exp
proximus
14.8%
±9.7
9
Muse Spark 1.2
proximus
12.0%
±5.8
10
Inkling
proximus
4.1%
±5.3
0%20%40%60%80%100%

Bar and value are mean@5, whiskers span worst@5–best@5. Green marks the top result. Cost and time are per-trial averages.

Analysis

Score for every model on every task; stronger color is higher. The "Best"column is the highest score any model achieved on that task. Click a cell for that pair's traces; click a Best cell to open the top-scoring run's trace.

0%100%