Skip to content
LLM Pulse

Leaderboard

Humanity's Last Exam

31 models scored · metric: accuracy. Prices are the cheapest listed across all serving providers. Pts/$ = score ÷ blended price (3×input + output) / 4 — how much score each dollar buys.

Best value · top 5 by pts per dollar
  1. 1.Step 3.7 Flash113
  2. 2.GPT-5.4 nano92.0
  3. 3.GPT-5.4 nano59.3
  4. 4.Nemotron 3 Ultra 550B A55B40.4
  5. 5.Muse Spark 1.131.1
#ModelScoreBest in /MBest out /MPts / $Context
🥇Claude Opus 5anthropic
64.7
$5$256.51M
🥈Claude Fable 5anthropic
64.5
$10$503.21M
🥉Muse Spark 1.1meta
62.1
$1.25$4.2531.11M
4Claude Fable 5anthropic
59
$10$503.01M
5GPT-5.4 Proopenai
58.7
$27$1601.01.05M
6Claude Opus 4.8anthropic
57.9
$5$255.81M
7GPT-5.5 Proopenai
57.2
$27.27$163.640.91.05M
8Claude Opus 5anthropic
56.3
$5$255.61M
9Claude Opus 4.7anthropic
54.7
$5$255.51M
10GLM-5.2zhipuai
54.7
$1.40$4.4025.41M
11GPT-5.5openai
52.2
$4.55$27.275.11.05M
12GPT-5.4openai
52.1
$2.20$1410.11.05M
13Claude Opus 4.8anthropic
49.8
$5$255.01M
14Step 3.7 Flashstepfun
47.2
$0.185$1.11113256K
15Claude Opus 4.7anthropic
46.9
$5$254.71M
16Claude Sonnet 4.6anthropic
46.8
$3$157.81M
17Gemini 3.1 Pro Previewgoogle
44.4
$2$129.91.05M
18GPT-5.5 Proopenai
43.1
$27.27$163.640.71.05M
19GPT-5.4 Proopenai
42.7
$27$1600.71.05M
20GPT-5.4 miniopenai
41.5
$0.68$427.5400K
21Qwen3.7 Maxalibaba
41.4
$2.50$7.5011.01M
22GPT-5.5openai
41.4
$4.55$27.274.01.05M
23GLM-5.2zhipuai
40.5
$1.40$4.4018.81M
24Gemini 3.5 Flashgoogle
40.2
$1.50$911.91.05M
25GPT-5.4openai
39.8
$2.20$147.71.05M
26GPT-5.4 nanoopenai
37.7
$0.18$1.1092.0400K
27Nemotron 3 Ultra 550B A55Bnvidia
37.4
$0.5$2.2040.41M
28Claude Sonnet 4.6anthropic
34.6
$3$155.81M
29GPT-5.4 miniopenai
28.2
$0.68$418.7400K
30Nemotron 3 Ultra 550B A55Bnvidia
26.7
$0.5$2.2028.91M
31GPT-5.4 nanoopenai
24.3
$0.18$1.1059.3400K

Cheapest scorer: GPT-5.4 nano at $0.18 input /M.