FrontierSWEV2

Benchmarking software engineering skill at the edge of human ability.

By

Leaderboard

Scores across all 34 tasks. Each model runs 5 trials per task with a 20-hour budget.

#ModelScore
1
GPT-6 Astra
proximus
65.5%
±8.9
2
Claude Opus 5.5
proximus
62.3%
±9.8
3
Gemini 4 Argon
proximus
55.0%
±9.9
4
GLM-5.3
proximus
30.2%
±11.5
5
Grok 4.7
proximus
29.5%
±12.6
6
Kimi K3
proximus
25.9%
±11.8
7
Muse Spark 1.3
proximus
25.8%
±14.2
8
Qwen3.8-Max-0902
proximus
17.8%
±9.4
9
DeepSeek V4 Flash Vision Exp
proximus
14.8%
±9.7
10
Inkling
proximus
4.1%
±5.3
0%20%40%60%80%100%

Bar and value are mean@5, whiskers span worst@5–best@5. Green marks the top result. Cost and time are per-trial averages.

*Gemini 4 Argon: Cache hit rate is much lower due to a non-production setting.

Analysis

Score for every model on every task; stronger color is higher. The "Best"column is the highest score any model achieved on that task. Click a cell for that pair's traces; click a Best cell to open the top-scoring run's trace.

0%100%