Reasoning

BIG-Bench Hard Leaderboard

BIG-Bench Hard (BBH) is the 23-task subset of BIG-Bench that earlier models found hardest — multi-step reasoning, logic, and word problems. It's a long-standing standard for tracking reasoning ability across model generations.

Source: epoch37 open models ranked+13 proprietaryData through Dec 2024

Open models ranked on BIG-Bench Hard

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 2DeepSeek v3 · 684.5B
87.5%
2 / 4Llama 3.1 405B · 405.9B
82.9%
3 / 6Qwen2.5 72B · 72.7B
79.8%
4 / 7Phi 3 Small 8k Instruct · 7.4B
79.1%
5 / 8DeepSeek v2 · 235.7B
78.8%
6 / 10Phi 3 Mini 4k Instruct · 3.8B
71.7%
7 / 11Yi 34B Chat · 34.4B
71.7%
8 / 12StableBeluga2 · 70B
69.3%
9 / 13Llama 2 70B HF · 69.0B
64.9%
10 / 15Phi 2 · 2.8B
59.4%
11 / 17Llama 2 70B Chat HF · 69.0B
58.5%
12 / 19Llama 2 13B Chat HF · 13.0B
58.2%
13 / 20Mistral 7B v0.1 · 7B
56.1%
14 / 21Gemma 7B · 8.5B
55.1%
15 / 22Qwen 14B Chat · 14.2B
55.0%
16 / 23Yi 34B · 34.4B
54.3%
17 / 24Qwen 14B · 14.2B
53.4%
18 / 25Internlm 20B · 20B
52.5%
19 / 27Baichuan2 13B Base · 13B
49.0%
20 / 28Baichuan2 13B Chat · 13B
47.2%
21 / 29Yi 6B Chat · 6.1B
47.2%
22 / 30Llama 2 13B HF · 13.0B
47.0%
23 / 31Qwen 7B · 7.7B
45.0%
24 / 34Baichuan 13B Base · 13B
43.0%
25 / 35Yi 6B · 6.1B
42.8%
26 / 36Internlm Chat 20B · 20B
42.4%
27 / 37Baichuan2 7B Base · 7B
41.6%
28 / 38Llama 2 7B HF · 6.7B
39.2%
29 / 41Falcon 40B · 41.8B
37.1%
30 / 42Internlm 7B · 7B
37.0%
31 / 44Gemma 2B · 2.5B
35.2%
32 / 45INTELLECT 1 Instruct · 10.2B
34.8%
33 / 46Chatglm2 6B · 6B
33.7%
34 / 47Llama 7B · 6.7B
33.5%
35 / 48Baichuan 7B · 7B
32.5%
36 / 49Falcon 7B · 7.2B
28.8%
37 / 50Qwen 1 8B · 1.8B
28.2%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

10B100Bmodel size (log scale) →87.5%28.2%DeepSeek v2 · 236B · 78.8%Yi 34B Chat · 34B · 71.7%StableBeluga2 · 70B · 69.3%Llama 2 70B HF · 69B · 64.9%Llama 2 70B Chat HF · 69B · 58.5%Llama 2 13B Chat HF · 13B · 58.2%Mistral 7B v0.1 · 7B · 56.1%Gemma 7B · 9B · 55.1%Qwen 14B Chat · 14B · 55.0%Yi 34B · 34B · 54.3%Qwen 14B · 14B · 53.4%Internlm 20B · 20B · 52.5%Baichuan2 13B Base · 13B · 49.0%Yi 6B Chat · 6B · 47.2%Baichuan2 13B Chat · 13B · 47.2%Llama 2 13B HF · 13B · 47.0%Qwen 7B · 8B · 45.0%Baichuan 13B Base · 13B · 43.0%Yi 6B · 6B · 42.8%Internlm Chat 20B · 20B · 42.4%Baichuan2 7B Base · 7B · 41.6%Llama 2 7B HF · 7B · 39.2%Falcon 40B · 42B · 37.1%Internlm 7B · 7B · 37.0%INTELLECT 1 Instruct · 10B · 34.8%Chatglm2 6B · 6B · 33.7%Llama 7B · 7B · 33.5%Baichuan 7B · 7B · 32.5%Falcon 7B · 7B · 28.8%Qwen 1 8B · 2B · 28.2%Qwen 1 8BGemma 2B · 3B · 35.2%Gemma 2BPhi 2 · 3B · 59.4%Phi 2Phi 3 Mini 4k Instruct · 4B · 71.7%Phi 3 Mini 4k InstructPhi 3 Small 8k Instruct · 7B · 79.1%Phi 3 Small 8k Instru…Qwen2.5 72B · 73B · 79.8%Qwen2.5 72BLlama 3.1 405B · 406B · 82.9%Llama 3.1 405BDeepSeek v3 · 685B · 87.5%DeepSeek v3
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Qwen 1 8B, 2B, score 28.2% — on the efficiency frontier (best score at its size or smaller).
  • Gemma 2B, 3B, score 35.2% — on the efficiency frontier (best score at its size or smaller).
  • Phi 2, 3B, score 59.4% — on the efficiency frontier (best score at its size or smaller).
  • Phi 3 Mini 4k Instruct, 4B, score 71.7% — on the efficiency frontier (best score at its size or smaller).
  • Phi 3 Small 8k Instruct, 7B, score 79.1% — on the efficiency frontier (best score at its size or smaller).
  • Qwen2.5 72B, 73B, score 79.8% — on the efficiency frontier (best score at its size or smaller).
  • Llama 3.1 405B, 406B, score 82.9% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek v3, 685B, score 87.5% — on the efficiency frontier (best score at its size or smaller).

BIG-Bench Hard: frequently asked questions

What is the best open LLM on BIG-Bench Hard?
DeepSeek v3 is the top open model on BIG-Bench Hard, scoring 87.5%. Among all models tested — including proprietary ones — it ranks #2. The top model overall is Gemini 1.5 Pro 001 (Google DeepMind) at 89.2%.
What's the best BIG-Bench Hard model you can run on a 24 GB GPU?
Phi 3 Small 8k Instruct is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 4 GB), scoring 79.1% on BIG-Bench Hard.
What's the best BIG-Bench Hard model you can run on a 12 GB GPU?
Phi 3 Small 8k Instruct is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 4 GB), scoring 79.1% on BIG-Bench Hard.
Can open models match proprietary models on BIG-Bench Hard?
Not quite on BIG-Bench Hard: the strongest proprietary model (Gemini 1.5 Pro 001) scores 89.2%, ahead of the best open model (DeepSeek v3) at 87.5% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.