Knowledge
TriviaQA Leaderboard
TriviaQA is a large set of trivia questions paired with evidence documents, testing a model's factual recall and reading comprehension across a wide range of everyday topics.
Source: epoch18 open models ranked+21 proprietaryData through Dec 2024
All models ranked on TriviaQA
Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.
| # | Model | Score |
|---|---|---|
| 1 | Llama 2 70B HF · 69.0B | 87.6% |
| 2 | Claude 2.0 · proprietary | 87.5% |
| 3 | Claude 1.3 · proprietary | 86.7% |
| 4 | PaLM 2 L · proprietary | 86.1% |
| 5 | Llama 65B · proprietary | 86.0% |
| 6 | GPT 3.5 Turbo (Nov 06) · proprietary | 85.8% |
| 7 | GPT 4 (Jun 13) · proprietary | 84.8% |
| 8 | Llama 2 34B · proprietary | 84.6% |
| 9 | Llama 33B · proprietary | 83.8% |
| 10 | DeepSeek v3 · 684.5B | 82.9% |
| 11 | Llama 3.1 405B · 405.9B | 82.7% |
| 12 | Mixtral 8x7B v0.1 · 46.7B | 82.2% |
| 13 | PaLM 2 M · proprietary | 81.7% |
| 14 | PaLM 540B · proprietary | 81.4% |
| 15 | DeepSeek v2 · 235.7B | 80.0% |
| 16 | Falcon 40B · 41.8B | 79.9% |
| 17 | Llama 2 13B HF · 13.0B | 79.6% |
| 18 | Claude Instant 1.1 · proprietary | 78.9% |
| 19 | Claude Instant 1.2 · proprietary | 78.7% |
| 20 | Llama 13B · proprietary | 77.9% |
| 21 | GLaM (MoE) · proprietary | 75.8% |
| 22 | Mistral 7B v0.1 · 7B | 75.2% |
| 23 | PaLM 2 S · proprietary | 75.2% |
| 24 | Phi 3 Medium 128K Instruct · proprietary | 73.9% |
| 25 | Llama 2 7B HF · 6.7B | 73.7% |
| 26 | Mpt 30B · proprietary | 73.6% |
| 27 | Gemma 7B · 8.5B | 72.3% |
| 28 | Qwen2.5 72B · 72.7B | 71.9% |
| 29 | Text Davinci 001 · proprietary | 71.2% |
| 30 | Llama 7B · 6.7B | 71.0% |
| 31 | Meta Llama 3 8B Instruct · 8.0B | 67.7% |
| 32 | Chinchilla (70B) · proprietary | 64.6% |
| 33 | Falcon 7B · 7.2B | 64.6% |
| 34 | Phi 3 Mini 4k Instruct · 3.8B | 64.0% |
| 35 | Mpt 7B · proprietary | 61.6% |
| 36 | Phi 3 Small 8k Instruct · 7.4B | 58.1% |
| 37 | Gopher (280B) · proprietary | 57.2% |
| 38 | Gemma 2B · 2.5B | 53.2% |
| 39 | Phi 2 · 2.8B | 45.2% |
Score vs model size
Which models give the most quality for their size — the ones worth running locally.
- Gemma 2B, 3B, score 53.2% — on the efficiency frontier (best score at its size or smaller).
- Phi 3 Mini 4k Instruct, 4B, score 64.0% — on the efficiency frontier (best score at its size or smaller).
- Llama 2 7B HF, 7B, score 73.7% — on the efficiency frontier (best score at its size or smaller).
- Mistral 7B v0.1, 7B, score 75.2% — on the efficiency frontier (best score at its size or smaller).
- Llama 2 13B HF, 13B, score 79.6% — on the efficiency frontier (best score at its size or smaller).
- Falcon 40B, 42B, score 79.9% — on the efficiency frontier (best score at its size or smaller).
- Mixtral 8x7B v0.1, 47B, score 82.2% — on the efficiency frontier (best score at its size or smaller).
- Llama 2 70B HF, 69B, score 87.6% — on the efficiency frontier (best score at its size or smaller).
TriviaQA: frequently asked questions
- What is the best open LLM on TriviaQA?
- Llama 2 70B HF is the top open model on TriviaQA, scoring 87.6%. Among all models tested — including proprietary ones — it ranks #1. That puts it ahead of every proprietary model we track, including Claude 2.0 (Anthropic) at 87.5%.
- What's the best TriviaQA model you can run on a 24 GB GPU?
- Falcon 40B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 23 GB), scoring 79.9% on TriviaQA.
- What's the best TriviaQA model you can run on a 12 GB GPU?
- Llama 2 13B HF is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 7 GB), scoring 79.6% on TriviaQA.
- Can open models match proprietary models on TriviaQA?
- Yes — the best open model (Llama 2 70B HF, 87.6%) matches or beats every proprietary model we track on TriviaQA.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.