Knowledge

ARC Challenge Leaderboard

ARC Challenge is a set of grade-school science questions, curated by the Allen Institute for AI (AI2) to exclude ones that simple word-matching or retrieval can answer. It measures science knowledge paired with the reasoning needed to apply it.

Source: epoch51 open models ranked+26 proprietaryData through Dec 2024

All models ranked on ARC Challenge

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1DeepSeek v3 · 684.5B
95.3%
2Llama 3.1 405B · 405.9B
95.3%
3Qwen2.5 72B · 72.7B
94.5%
4DeepSeek v2 · 235.7B
92.2%
5Phi 3 Medium 128K Instruct · proprietary
91.6%
6Phi 3 Small 8k Instruct · 7.4B
90.7%
7GPT 3.5 Turbo (Nov 06) · proprietary
87.4%
8Mixtral 8x7B v0.1 · 46.7B
87.3%
9Claude Instant 1.2 · proprietary
86.3%
10StableBeluga2 · 70B
86.1%
11Claude Instant 1.1 · proprietary
85.7%
12PaLM 540B · proprietary
85.2%
13Text Davinci 002 · proprietary
85.2%
14Phi 3 Mini 4k Instruct · 3.8B
84.9%
15Qwen 14B · 14.2B
84.4%
16Meta Llama 3 8B Instruct · 8.0B
82.8%
17Internlm 20B · 20B
81.7%
18Mistral 7B v0.1 · 7B
78.6%
19Gemma 7B · 8.5B
78.3%
20Llama 2 70B HF · 69.0B
78.3%
21Phi 2 · 2.8B
75.9%
22Qwen 7B · 7.7B
75.3%
23Qwen2.5 Coder 32B · 32.8B
70.5%
24Internlm 7B · 7B
69.5%
25Llama 65B · proprietary
69.5%
26PaLM 2 L · proprietary
69.2%
27Falcon 180B · 180B
67.8%
28Llama 33B · proprietary
67.5%
29Qwen2.5 Coder 14B · 14.8B
66.0%
30PaLM 2 M · proprietary
64.9%
31DeepSeek Coder V2 Base · proprietary
64.3%
32Falcon 40B · 41.8B
61.9%
33Chatglm2 6B · 6B
61.0%
34Qwen2.5 Coder 7B · 7.6B
60.9%
35Llama 2 13B HF · 13.0B
60.3%
36PaLM 2 S · proprietary
59.6%
37DeepSeek Coder v2 Lite Base · 15.7B
57.3%
38Yi 9B · 8.8B
55.6%
39Nemotron 4 15B · proprietary
55.5%
40INTELLECT 1 Instruct · 10.2B
54.5%
41Llama 2 34B · proprietary
54.5%
42Qwen 1 8B · 1.8B
53.2%
43Text Davinci 001 · proprietary
53.2%
44Qwen2.5 Coder 3B · 3.1B
52.9%
45Llama 13B · proprietary
52.7%
46PaLM 62B · proprietary
52.5%
47Mpt 30B · proprietary
50.6%
48Yi 6B · 6.1B
50.3%
49Falcon 7B · 7.2B
47.9%
50Llama 7B · 6.7B
47.6%
51Starcoder2 15B · 16.0B
47.2%
52Llama 2 7B HF · 6.7B
45.9%
53Qwen2.5 Coder 1.5B · 1.5B
45.2%
54Phi 1 5 · 1.4B
44.4%
55Vicuna 13B v1.1 · proprietary
43.2%
56Mpt 7B · proprietary
42.6%
57DeepSeek Coder 33B Base · proprietary
42.2%
58Gemma 2B · 2.5B
42.1%
59Xgen 7B 8k Base · 7B
41.2%
60GPT Neox 20B · 20.7B
41.1%
61Dolly v2 12B · proprietary
39.6%
62RedPajama INCITE 7B Base · proprietary
39.1%
63Open Llama 7B · proprietary
38.7%
64Starcoder2 7B · 7.2B
38.7%
65Baichuan2 13B Base · 13B
38.0%
66Deepseek Coder 6.7B Base · 6.7B
36.4%
67GPT J 6B · 6B
36.3%
68Opt 13B · proprietary
35.8%
69CodeQwen1.5 7B · 7.3B
35.7%
70Qwen2.5 Coder 0.5B · 494M
34.4%
71Starcoder2 3B · 3.0B
34.2%
72Baichuan2 7B Base · 7B
32.5%
73Cerebras GPT 13B · 13B
32.4%
74Stablelm Tuned Alpha 7B · 7B
27.0%
75Deepseek Coder 1.3B Base · 1.3B
25.4%
76Gpt2 Xl · 1.6B
25.0%
77Opt 1.3b · proprietary
23.2%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

1B10B100Bmodel size (log scale) →95.3%25.0%DeepSeek v3 · 685B · 95.3%DeepSeek v2 · 236B · 92.2%Mixtral 8x7B v0.1 · 47B · 87.3%StableBeluga2 · 70B · 86.1%Qwen 14B · 14B · 84.4%Meta Llama 3 8B Instruct · 8B · 82.8%Internlm 20B · 20B · 81.7%Mistral 7B v0.1 · 7B · 78.6%Gemma 7B · 9B · 78.3%Llama 2 70B HF · 69B · 78.3%Qwen 7B · 8B · 75.3%Qwen2.5 Coder 32B · 33B · 70.5%Internlm 7B · 7B · 69.5%Falcon 180B · 180B · 67.8%Qwen2.5 Coder 14B · 15B · 66.0%Falcon 40B · 42B · 61.9%Chatglm2 6B · 6B · 61.0%Qwen2.5 Coder 7B · 8B · 60.9%Llama 2 13B HF · 13B · 60.3%DeepSeek Coder v2 Lite Base · 16B · 57.3%Yi 9B · 9B · 55.6%INTELLECT 1 Instruct · 10B · 54.5%Qwen2.5 Coder 3B · 3B · 52.9%Yi 6B · 6B · 50.3%Falcon 7B · 7B · 47.9%Llama 7B · 7B · 47.6%Starcoder2 15B · 16B · 47.2%Llama 2 7B HF · 7B · 45.9%Gemma 2B · 3B · 42.1%Xgen 7B 8k Base · 7B · 41.2%GPT Neox 20B · 21B · 41.1%Starcoder2 7B · 7B · 38.7%Baichuan2 13B Base · 13B · 38.0%Deepseek Coder 6.7B Base · 7B · 36.4%GPT J 6B · 6B · 36.3%CodeQwen1.5 7B · 7B · 35.7%Starcoder2 3B · 3B · 34.2%Baichuan2 7B Base · 7B · 32.5%Cerebras GPT 13B · 13B · 32.4%Stablelm Tuned Alpha 7B · 7B · 27.0%Deepseek Coder 1.3B Base · 1B · 25.4%Gpt2 Xl · 2B · 25.0%Qwen2.5 Coder 0.5B · 494M · 34.4%Qwen2.5 Coder 0.5BPhi 1 5 · 1B · 44.4%Phi 1 5Qwen2.5 Coder 1.5B · 2B · 45.2%Qwen2.5 Coder 1.5BQwen 1 8B · 2B · 53.2%Qwen 1 8BPhi 2 · 3B · 75.9%Phi 2Phi 3 Mini 4k Instruct · 4B · 84.9%Phi 3 Mini 4k InstructPhi 3 Small 8k Instruct · 7B · 90.7%Phi 3 Small 8k Instru…Qwen2.5 72B · 73B · 94.5%Qwen2.5 72BLlama 3.1 405B · 406B · 95.3%Llama 3.1 405B
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Qwen2.5 Coder 0.5B, 494M, score 34.4% — on the efficiency frontier (best score at its size or smaller).
  • Phi 1 5, 1B, score 44.4% — on the efficiency frontier (best score at its size or smaller).
  • Qwen2.5 Coder 1.5B, 2B, score 45.2% — on the efficiency frontier (best score at its size or smaller).
  • Qwen 1 8B, 2B, score 53.2% — on the efficiency frontier (best score at its size or smaller).
  • Phi 2, 3B, score 75.9% — on the efficiency frontier (best score at its size or smaller).
  • Phi 3 Mini 4k Instruct, 4B, score 84.9% — on the efficiency frontier (best score at its size or smaller).
  • Phi 3 Small 8k Instruct, 7B, score 90.7% — on the efficiency frontier (best score at its size or smaller).
  • Qwen2.5 72B, 73B, score 94.5% — on the efficiency frontier (best score at its size or smaller).
  • Llama 3.1 405B, 406B, score 95.3% — on the efficiency frontier (best score at its size or smaller).

ARC Challenge: frequently asked questions

What is the best open LLM on ARC Challenge?
DeepSeek v3 is the top open model on ARC Challenge, scoring 95.3%. Among all models tested — including proprietary ones — it ranks #1. That puts it ahead of every proprietary model we track, including Phi 3 Medium 128K Instruct (Microsoft) at 91.6%.
What's the best ARC Challenge model you can run on a 24 GB GPU?
Phi 3 Small 8k Instruct is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 4 GB), scoring 90.7% on ARC Challenge.
What's the best ARC Challenge model you can run on a 12 GB GPU?
Phi 3 Small 8k Instruct is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 4 GB), scoring 90.7% on ARC Challenge.
Can open models match proprietary models on ARC Challenge?
Yes — the best open model (DeepSeek v3, 95.3%) matches or beats every proprietary model we track on ARC Challenge.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.