Skip to content
LLM Pulse

Leaderboard

SWE-Atlas Refactoring

11 models scored · metric: score. Prices are the cheapest listed across all serving providers. Pts/$ = score ÷ blended price (3×input + output) / 4 — how much score each dollar buys.

Best value · top 5 by pts per dollar
  1. 1.MiniMax-M2.537.2
  2. 2.Kimi K2.529.9
  3. 3.GLM-515.6
  4. 4.GPT-5.3 Codex9.5
  5. 5.Gemini 3 Flash Preview8.9
#ModelScoreBest in /MBest out /MPts / $Context
🥇Claude Opus 4.7anthropic
48.57
$5$254.91M
🥈GPT-5.5openai
44.79
$4.55$27.274.41.05M
🥉GPT-5.4openai
44.29
$2.20$148.61.05M
4GPT-5.3 Codexopenai
42.38
$1.60$139.5400K
5Claude Opus 4.6anthropic
35.58
$5$253.61M
6Gemini 3.1 Pro Previewgoogle
33.81
$2$127.51.05M
7Claude Sonnet 4.6anthropic
32.21
$3$155.41M
8GLM-5zhipuai
24.24
$1$3.2015.6205K
9Kimi K2.5moonshotai
20.95
$0.3$1.9029.9262K
10MiniMax-M2.5minimax
19.52
$0.3$1.2037.2205K
11Gemini 3 Flash Previewgoogle
10
$0.5$38.91.05M

Cheapest scorer: Kimi K2.5 at $0.3 input /M.