Knowledge

ARC Challenge Leaderboard

ARC Challenge is a set of grade-school science questions, curated by the Allen Institute for AI (AI2) to exclude ones that simple word-matching or retrieval can answer. It measures science knowledge paired with the reasoning needed to apply it.

Source: epoch51 open models ranked+26 proprietaryData through Dec 2024

Open models ranked on ARC Challenge

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 1DeepSeek v3 · 684.5B
95.3%
2 / 2Llama 3.1 405B · 405.9B
95.3%
3 / 3Qwen2.5 72B · 72.7B
94.5%
4 / 4DeepSeek v2 · 235.7B
92.2%
5 / 6Phi 3 Small 8k Instruct · 7.4B
90.7%
6 / 8Mixtral 8x7B v0.1 · 46.7B
87.3%
7 / 10StableBeluga2 · 70B
86.1%
8 / 14Phi 3 Mini 4k Instruct · 3.8B
84.9%
9 / 15Qwen 14B · 14.2B
84.4%
10 / 16Meta Llama 3 8B Instruct · 8.0B
82.8%
11 / 17Internlm 20B · 20B
81.7%
12 / 18Mistral 7B v0.1 · 7B
78.6%
13 / 19Gemma 7B · 8.5B
78.3%
14 / 20Llama 2 70B HF · 69.0B
78.3%
15 / 21Phi 2 · 2.8B
75.9%
16 / 22Qwen 7B · 7.7B
75.3%
17 / 23Qwen2.5 Coder 32B · 32.8B
70.5%
18 / 24Internlm 7B · 7B
69.5%
19 / 27Falcon 180B · 180B
67.8%
20 / 29Qwen2.5 Coder 14B · 14.8B
66.0%
21 / 32Falcon 40B · 41.8B
61.9%
22 / 33Chatglm2 6B · 6B
61.0%
23 / 34Qwen2.5 Coder 7B · 7.6B
60.9%
24 / 35Llama 2 13B HF · 13.0B
60.3%
25 / 37DeepSeek Coder v2 Lite Base · 15.7B
57.3%
26 / 38Yi 9B · 8.8B
55.6%
27 / 40INTELLECT 1 Instruct · 10.2B
54.5%
28 / 42Qwen 1 8B · 1.8B
53.2%
29 / 44Qwen2.5 Coder 3B · 3.1B
52.9%
30 / 48Yi 6B · 6.1B
50.3%
31 / 49Falcon 7B · 7.2B
47.9%
32 / 50Llama 7B · 6.7B
47.6%
33 / 51Starcoder2 15B · 16.0B
47.2%
34 / 52Llama 2 7B HF · 6.7B
45.9%
35 / 53Qwen2.5 Coder 1.5B · 1.5B
45.2%
36 / 54Phi 1 5 · 1.4B
44.4%
37 / 58Gemma 2B · 2.5B
42.1%
38 / 59Xgen 7B 8k Base · 7B
41.2%
39 / 60GPT Neox 20B · 20.7B
41.1%
40 / 64Starcoder2 7B · 7.2B
38.7%
41 / 65Baichuan2 13B Base · 13B
38.0%
42 / 66Deepseek Coder 6.7B Base · 6.7B
36.4%
43 / 67GPT J 6B · 6B
36.3%
44 / 69CodeQwen1.5 7B · 7.3B
35.7%
45 / 70Qwen2.5 Coder 0.5B · 494M
34.4%
46 / 71Starcoder2 3B · 3.0B
34.2%
47 / 72Baichuan2 7B Base · 7B
32.5%
48 / 73Cerebras GPT 13B · 13B
32.4%
49 / 74Stablelm Tuned Alpha 7B · 7B
27.0%
50 / 75Deepseek Coder 1.3B Base · 1.3B
25.4%
51 / 76Gpt2 Xl · 1.6B
25.0%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

1B10B100Bmodel size (log scale) →95.3%25.0%DeepSeek v3 · 685B · 95.3%DeepSeek v2 · 236B · 92.2%Mixtral 8x7B v0.1 · 47B · 87.3%StableBeluga2 · 70B · 86.1%Qwen 14B · 14B · 84.4%Meta Llama 3 8B Instruct · 8B · 82.8%Internlm 20B · 20B · 81.7%Mistral 7B v0.1 · 7B · 78.6%Gemma 7B · 9B · 78.3%Llama 2 70B HF · 69B · 78.3%Qwen 7B · 8B · 75.3%Qwen2.5 Coder 32B · 33B · 70.5%Internlm 7B · 7B · 69.5%Falcon 180B · 180B · 67.8%Qwen2.5 Coder 14B · 15B · 66.0%Falcon 40B · 42B · 61.9%Chatglm2 6B · 6B · 61.0%Qwen2.5 Coder 7B · 8B · 60.9%Llama 2 13B HF · 13B · 60.3%DeepSeek Coder v2 Lite Base · 16B · 57.3%Yi 9B · 9B · 55.6%INTELLECT 1 Instruct · 10B · 54.5%Qwen2.5 Coder 3B · 3B · 52.9%Yi 6B · 6B · 50.3%Falcon 7B · 7B · 47.9%Llama 7B · 7B · 47.6%Starcoder2 15B · 16B · 47.2%Llama 2 7B HF · 7B · 45.9%Gemma 2B · 3B · 42.1%Xgen 7B 8k Base · 7B · 41.2%GPT Neox 20B · 21B · 41.1%Starcoder2 7B · 7B · 38.7%Baichuan2 13B Base · 13B · 38.0%Deepseek Coder 6.7B Base · 7B · 36.4%GPT J 6B · 6B · 36.3%CodeQwen1.5 7B · 7B · 35.7%Starcoder2 3B · 3B · 34.2%Baichuan2 7B Base · 7B · 32.5%Cerebras GPT 13B · 13B · 32.4%Stablelm Tuned Alpha 7B · 7B · 27.0%Deepseek Coder 1.3B Base · 1B · 25.4%Gpt2 Xl · 2B · 25.0%Qwen2.5 Coder 0.5B · 494M · 34.4%Qwen2.5 Coder 0.5BPhi 1 5 · 1B · 44.4%Phi 1 5Qwen2.5 Coder 1.5B · 2B · 45.2%Qwen2.5 Coder 1.5BQwen 1 8B · 2B · 53.2%Qwen 1 8BPhi 2 · 3B · 75.9%Phi 2Phi 3 Mini 4k Instruct · 4B · 84.9%Phi 3 Mini 4k InstructPhi 3 Small 8k Instruct · 7B · 90.7%Phi 3 Small 8k Instru…Qwen2.5 72B · 73B · 94.5%Qwen2.5 72BLlama 3.1 405B · 406B · 95.3%Llama 3.1 405B
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Qwen2.5 Coder 0.5B, 494M, score 34.4% — on the efficiency frontier (best score at its size or smaller).
  • Phi 1 5, 1B, score 44.4% — on the efficiency frontier (best score at its size or smaller).
  • Qwen2.5 Coder 1.5B, 2B, score 45.2% — on the efficiency frontier (best score at its size or smaller).
  • Qwen 1 8B, 2B, score 53.2% — on the efficiency frontier (best score at its size or smaller).
  • Phi 2, 3B, score 75.9% — on the efficiency frontier (best score at its size or smaller).
  • Phi 3 Mini 4k Instruct, 4B, score 84.9% — on the efficiency frontier (best score at its size or smaller).
  • Phi 3 Small 8k Instruct, 7B, score 90.7% — on the efficiency frontier (best score at its size or smaller).
  • Qwen2.5 72B, 73B, score 94.5% — on the efficiency frontier (best score at its size or smaller).
  • Llama 3.1 405B, 406B, score 95.3% — on the efficiency frontier (best score at its size or smaller).

ARC Challenge: frequently asked questions

What is the best open LLM on ARC Challenge?
DeepSeek v3 is the top open model on ARC Challenge, scoring 95.3%. Among all models tested — including proprietary ones — it ranks #1. That puts it ahead of every proprietary model we track, including Phi 3 Medium 128K Instruct (Microsoft) at 91.6%.
What's the best ARC Challenge model you can run on a 24 GB GPU?
Phi 3 Small 8k Instruct is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 4 GB), scoring 90.7% on ARC Challenge.
What's the best ARC Challenge model you can run on a 12 GB GPU?
Phi 3 Small 8k Instruct is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 4 GB), scoring 90.7% on ARC Challenge.
Can open models match proprietary models on ARC Challenge?
Yes — the best open model (DeepSeek v3, 95.3%) matches or beats every proprietary model we track on ARC Challenge.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.