Knowledge
ARC Challenge Leaderboard
ARC Challenge is a set of grade-school science questions, curated by the Allen Institute for AI (AI2) to exclude ones that simple word-matching or retrieval can answer. It measures science knowledge paired with the reasoning needed to apply it.
Source: epoch51 open models ranked+26 proprietaryData through Dec 2024
All models ranked on ARC Challenge
Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek v3 · 684.5B | 95.3% |
| 2 | Llama 3.1 405B · 405.9B | 95.3% |
| 3 | Qwen2.5 72B · 72.7B | 94.5% |
| 4 | DeepSeek v2 · 235.7B | 92.2% |
| 5 | Phi 3 Medium 128K Instruct · proprietary | 91.6% |
| 6 | Phi 3 Small 8k Instruct · 7.4B | 90.7% |
| 7 | GPT 3.5 Turbo (Nov 06) · proprietary | 87.4% |
| 8 | Mixtral 8x7B v0.1 · 46.7B | 87.3% |
| 9 | Claude Instant 1.2 · proprietary | 86.3% |
| 10 | StableBeluga2 · 70B | 86.1% |
| 11 | Claude Instant 1.1 · proprietary | 85.7% |
| 12 | PaLM 540B · proprietary | 85.2% |
| 13 | Text Davinci 002 · proprietary | 85.2% |
| 14 | Phi 3 Mini 4k Instruct · 3.8B | 84.9% |
| 15 | Qwen 14B · 14.2B | 84.4% |
| 16 | Meta Llama 3 8B Instruct · 8.0B | 82.8% |
| 17 | Internlm 20B · 20B | 81.7% |
| 18 | Mistral 7B v0.1 · 7B | 78.6% |
| 19 | Gemma 7B · 8.5B | 78.3% |
| 20 | Llama 2 70B HF · 69.0B | 78.3% |
| 21 | Phi 2 · 2.8B | 75.9% |
| 22 | Qwen 7B · 7.7B | 75.3% |
| 23 | Qwen2.5 Coder 32B · 32.8B | 70.5% |
| 24 | Internlm 7B · 7B | 69.5% |
| 25 | Llama 65B · proprietary | 69.5% |
| 26 | PaLM 2 L · proprietary | 69.2% |
| 27 | Falcon 180B · 180B | 67.8% |
| 28 | Llama 33B · proprietary | 67.5% |
| 29 | Qwen2.5 Coder 14B · 14.8B | 66.0% |
| 30 | PaLM 2 M · proprietary | 64.9% |
| 31 | DeepSeek Coder V2 Base · proprietary | 64.3% |
| 32 | Falcon 40B · 41.8B | 61.9% |
| 33 | Chatglm2 6B · 6B | 61.0% |
| 34 | Qwen2.5 Coder 7B · 7.6B | 60.9% |
| 35 | Llama 2 13B HF · 13.0B | 60.3% |
| 36 | PaLM 2 S · proprietary | 59.6% |
| 37 | DeepSeek Coder v2 Lite Base · 15.7B | 57.3% |
| 38 | Yi 9B · 8.8B | 55.6% |
| 39 | Nemotron 4 15B · proprietary | 55.5% |
| 40 | INTELLECT 1 Instruct · 10.2B | 54.5% |
| 41 | Llama 2 34B · proprietary | 54.5% |
| 42 | Qwen 1 8B · 1.8B | 53.2% |
| 43 | Text Davinci 001 · proprietary | 53.2% |
| 44 | Qwen2.5 Coder 3B · 3.1B | 52.9% |
| 45 | Llama 13B · proprietary | 52.7% |
| 46 | PaLM 62B · proprietary | 52.5% |
| 47 | Mpt 30B · proprietary | 50.6% |
| 48 | Yi 6B · 6.1B | 50.3% |
| 49 | Falcon 7B · 7.2B | 47.9% |
| 50 | Llama 7B · 6.7B | 47.6% |
| 51 | Starcoder2 15B · 16.0B | 47.2% |
| 52 | Llama 2 7B HF · 6.7B | 45.9% |
| 53 | Qwen2.5 Coder 1.5B · 1.5B | 45.2% |
| 54 | Phi 1 5 · 1.4B | 44.4% |
| 55 | Vicuna 13B v1.1 · proprietary | 43.2% |
| 56 | Mpt 7B · proprietary | 42.6% |
| 57 | DeepSeek Coder 33B Base · proprietary | 42.2% |
| 58 | Gemma 2B · 2.5B | 42.1% |
| 59 | Xgen 7B 8k Base · 7B | 41.2% |
| 60 | GPT Neox 20B · 20.7B | 41.1% |
| 61 | Dolly v2 12B · proprietary | 39.6% |
| 62 | RedPajama INCITE 7B Base · proprietary | 39.1% |
| 63 | Open Llama 7B · proprietary | 38.7% |
| 64 | Starcoder2 7B · 7.2B | 38.7% |
| 65 | Baichuan2 13B Base · 13B | 38.0% |
| 66 | Deepseek Coder 6.7B Base · 6.7B | 36.4% |
| 67 | GPT J 6B · 6B | 36.3% |
| 68 | Opt 13B · proprietary | 35.8% |
| 69 | CodeQwen1.5 7B · 7.3B | 35.7% |
| 70 | Qwen2.5 Coder 0.5B · 494M | 34.4% |
| 71 | Starcoder2 3B · 3.0B | 34.2% |
| 72 | Baichuan2 7B Base · 7B | 32.5% |
| 73 | Cerebras GPT 13B · 13B | 32.4% |
| 74 | Stablelm Tuned Alpha 7B · 7B | 27.0% |
| 75 | Deepseek Coder 1.3B Base · 1.3B | 25.4% |
| 76 | Gpt2 Xl · 1.6B | 25.0% |
| 77 | Opt 1.3b · proprietary | 23.2% |
Score vs model size
Which models give the most quality for their size — the ones worth running locally.
- Qwen2.5 Coder 0.5B, 494M, score 34.4% — on the efficiency frontier (best score at its size or smaller).
- Phi 1 5, 1B, score 44.4% — on the efficiency frontier (best score at its size or smaller).
- Qwen2.5 Coder 1.5B, 2B, score 45.2% — on the efficiency frontier (best score at its size or smaller).
- Qwen 1 8B, 2B, score 53.2% — on the efficiency frontier (best score at its size or smaller).
- Phi 2, 3B, score 75.9% — on the efficiency frontier (best score at its size or smaller).
- Phi 3 Mini 4k Instruct, 4B, score 84.9% — on the efficiency frontier (best score at its size or smaller).
- Phi 3 Small 8k Instruct, 7B, score 90.7% — on the efficiency frontier (best score at its size or smaller).
- Qwen2.5 72B, 73B, score 94.5% — on the efficiency frontier (best score at its size or smaller).
- Llama 3.1 405B, 406B, score 95.3% — on the efficiency frontier (best score at its size or smaller).
ARC Challenge: frequently asked questions
- What is the best open LLM on ARC Challenge?
- DeepSeek v3 is the top open model on ARC Challenge, scoring 95.3%. Among all models tested — including proprietary ones — it ranks #1. That puts it ahead of every proprietary model we track, including Phi 3 Medium 128K Instruct (Microsoft) at 91.6%.
- What's the best ARC Challenge model you can run on a 24 GB GPU?
- Phi 3 Small 8k Instruct is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 4 GB), scoring 90.7% on ARC Challenge.
- What's the best ARC Challenge model you can run on a 12 GB GPU?
- Phi 3 Small 8k Instruct is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 4 GB), scoring 90.7% on ARC Challenge.
- Can open models match proprietary models on ARC Challenge?
- Yes — the best open model (DeepSeek v3, 95.3%) matches or beats every proprietary model we track on ARC Challenge.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.