Reasoning

LiveBench Reasoning Leaderboard

LiveBench Reasoning measures logical, multi-step reasoning using contamination-free questions that are refreshed regularly, so models cannot have trained on the test set.

Source: livebench16 open models ranked+38 proprietaryData through Jun 2026

All models ranked on LiveBench Reasoning

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1GPT 6 Astra Max · proprietary
92.7
2Claude Fable 5.1 Max · proprietary
91.7
3GPT 5.6 Sol Max · proprietary
91.7
4Claude Opus 5 Max · proprietary
91.2
5Kimi K3 · 2779.9B
90.7
6GPT 5.6 Terra Max · proprietary
90.6
7Grok 4.6 · proprietary
90.5
8Smaug Agentic · proprietary
90.3
9Muse Spark 1.2 · proprietary
90.0
10Claude Fable 5 Max · proprietary
89.7
11GPT 5.5 · proprietary
89.7
12Muse Spark 1.3 · proprietary
89.7
13Gemini 3.8 Flash · proprietary
89.3
14Claude Opus 4.8 Max · proprietary
89.2
15Claude Sonnet 5 · proprietary
88.7
16Claude Opus 4.6 · proprietary
88.7
17Qwen3.8 Max · proprietary
88.2
18GPT 5.4 · proprietary
88.1
19Gemini 3.7 Flash · proprietary
87.8
20Muse Spark 1.1 · proprietary
87.7
21Qwen3.8 Flash Next · 180.0B
87.4
22Claude Opus 4.7 · proprietary
87.2
23Grok 4.5 · proprietary
87.2
24DeepSeek V4 Flash 0731 · 304.2B
86.6
25DeepSeek V4 Pro 0813 · 1650.5B
85.8
26GLM 5.3 · 753.3B
85.8
27GPT 5.6 Luna Max · proprietary
85.6
28DeepSeek v4 Flash Vision · proprietary
85.4
29Gemini 3.6 Flash · proprietary
85.2
30Claude Sonnet 4.6 · proprietary
84.8
31Gemini 3.1 Pro · proprietary
84.0
32Qwen3.7 Max · proprietary
83.3
33GPT 5.2 · proprietary
83.2
34Kimi K2.7 Code · 1026.9B
82.8
35DeepSeek V4 Pro · 1598.8B
82.7
36Gemini 3.5 Flash · proprietary
82.0
37GPT 5.4 Nano · proprietary
81.1
38Claude Opus 4.5 · proprietary
80.1
39Qwen3.8 27B · 27.8B
80.0
40Kimi K2.6 · 1026.9B
79.4
41GLM 5.2 · 753.3B
78.6
42Inkling · 952.4B
78.3
43GPT 5.2 Codex · proprietary
77.7
44GLM 5.3 Flash · 321.3B
77.6
45Ox Alpha Max · proprietary
76.6
46Grok Build 0.1 · proprietary
76.4
47Qwen3.6 Plus · proprietary
75.8
48NVIDIA Nemotron 3 Ultra 550B A55B BF16 · 560.5B
74.7
49MiniMax M3 · 427.0B
74.5
50GPT 5.4 Mini · proprietary
71.3
51Grok 4.3 · proprietary
70.8
52DeepSeek V4 Flash · 290.9B
70.6
53Qwen3.6 27B · 27.8B
70.3
54Gemini 3.5 Flash Lite · proprietary
60.2

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

100B1Tmodel size (log scale) →90.770.3DeepSeek V4 Flash 0731 · 304B · 86.6DeepSeek V4 Pro 0813 · 1.7T · 85.8GLM 5.3 · 753B · 85.8Kimi K2.7 Code · 1T · 82.8DeepSeek V4 Pro · 1.6T · 82.7Kimi K2.6 · 1T · 79.4GLM 5.2 · 753B · 78.6Inkling · 952B · 78.3GLM 5.3 Flash · 321B · 77.6NVIDIA Nemotron 3 Ultra 550B A55B BF16 · 561B · 74.7MiniMax M3 · 427B · 74.5DeepSeek V4 Flash · 291B · 70.6Qwen3.6 27B · 28B · 70.3Qwen3.8 27B · 28B · 80.0Qwen3.8 27BQwen3.8 Flash Next · 180B · 87.4Qwen3.8 Flash NextKimi K3 · 2.8T · 90.7Kimi K3
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Qwen3.8 27B, 28B, score 80.0 — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.8 Flash Next, 180B, score 87.4 — on the efficiency frontier (best score at its size or smaller).
  • Kimi K3, 2.8T, score 90.7 — on the efficiency frontier (best score at its size or smaller).

LiveBench Reasoning: frequently asked questions

What is the best open LLM on LiveBench Reasoning?
Kimi K3 is the top open model on LiveBench Reasoning, scoring 90.7. Among all models tested — including proprietary ones — it ranks #5. The top model overall is GPT 6 Astra Max (OpenAI) at 92.7.
What's the best LiveBench Reasoning model you can run on a 24 GB GPU?
Qwen3.8 27B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 15 GB), scoring 80.0 on LiveBench Reasoning.
Can open models match proprietary models on LiveBench Reasoning?
Not quite on LiveBench Reasoning: the strongest proprietary model (GPT 6 Astra Max) scores 92.7, ahead of the best open model (Kimi K3) at 90.7 — but you can run the open one yourself.

Scores aggregated from livebench. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.