Reasoning

LiveBench Reasoning Leaderboard

LiveBench Reasoning measures logical, multi-step reasoning using contamination-free questions that are refreshed regularly, so models cannot have trained on the test set.

Source: livebench16 open models ranked+38 proprietaryData through Jun 2026

Open models ranked on LiveBench Reasoning

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 5Kimi K3 · 2779.9B
90.7
2 / 21Qwen3.8 Flash Next · 180.0B
87.4
3 / 24DeepSeek V4 Flash 0731 · 304.2B
86.6
4 / 25DeepSeek V4 Pro 0813 · 1650.5B
85.8
5 / 26GLM 5.3 · 753.3B
85.8
6 / 34Kimi K2.7 Code · 1026.9B
82.8
7 / 35DeepSeek V4 Pro · 1598.8B
82.7
8 / 39Qwen3.8 27B · 27.8B
80.0
9 / 40Kimi K2.6 · 1026.9B
79.4
10 / 41GLM 5.2 · 753.3B
78.6
11 / 42Inkling · 952.4B
78.3
12 / 44GLM 5.3 Flash · 321.3B
77.6
13 / 48NVIDIA Nemotron 3 Ultra 550B A55B BF16 · 560.5B
74.7
14 / 49MiniMax M3 · 427.0B
74.5
15 / 52DeepSeek V4 Flash · 290.9B
70.6
16 / 53Qwen3.6 27B · 27.8B
70.3

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

100B1Tmodel size (log scale) →90.770.3DeepSeek V4 Flash 0731 · 304B · 86.6DeepSeek V4 Pro 0813 · 1.7T · 85.8GLM 5.3 · 753B · 85.8Kimi K2.7 Code · 1T · 82.8DeepSeek V4 Pro · 1.6T · 82.7Kimi K2.6 · 1T · 79.4GLM 5.2 · 753B · 78.6Inkling · 952B · 78.3GLM 5.3 Flash · 321B · 77.6NVIDIA Nemotron 3 Ultra 550B A55B BF16 · 561B · 74.7MiniMax M3 · 427B · 74.5DeepSeek V4 Flash · 291B · 70.6Qwen3.6 27B · 28B · 70.3Qwen3.8 27B · 28B · 80.0Qwen3.8 27BQwen3.8 Flash Next · 180B · 87.4Qwen3.8 Flash NextKimi K3 · 2.8T · 90.7Kimi K3
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Qwen3.8 27B, 28B, score 80.0 — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.8 Flash Next, 180B, score 87.4 — on the efficiency frontier (best score at its size or smaller).
  • Kimi K3, 2.8T, score 90.7 — on the efficiency frontier (best score at its size or smaller).

LiveBench Reasoning: frequently asked questions

What is the best open LLM on LiveBench Reasoning?
Kimi K3 is the top open model on LiveBench Reasoning, scoring 90.7. Among all models tested — including proprietary ones — it ranks #5. The top model overall is GPT 6 Astra Max (OpenAI) at 92.7.
What's the best LiveBench Reasoning model you can run on a 24 GB GPU?
Qwen3.8 27B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 15 GB), scoring 80.0 on LiveBench Reasoning.
Can open models match proprietary models on LiveBench Reasoning?
Not quite on LiveBench Reasoning: the strongest proprietary model (GPT 6 Astra Max) scores 92.7, ahead of the best open model (Kimi K3) at 90.7 — but you can run the open one yourself.

Scores aggregated from livebench. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.