Knowledge

TriviaQA Leaderboard

TriviaQA is a large set of trivia questions paired with evidence documents, testing a model's factual recall and reading comprehension across a wide range of everyday topics.

Source: epoch18 open models ranked+21 proprietaryData through Dec 2024

All models ranked on TriviaQA

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1Llama 2 70B HF · 69.0B
87.6%
2Claude 2.0 · proprietary
87.5%
3Claude 1.3 · proprietary
86.7%
4PaLM 2 L · proprietary
86.1%
5Llama 65B · proprietary
86.0%
6GPT 3.5 Turbo (Nov 06) · proprietary
85.8%
7GPT 4 (Jun 13) · proprietary
84.8%
8Llama 2 34B · proprietary
84.6%
9Llama 33B · proprietary
83.8%
10DeepSeek v3 · 684.5B
82.9%
11Llama 3.1 405B · 405.9B
82.7%
12Mixtral 8x7B v0.1 · 46.7B
82.2%
13PaLM 2 M · proprietary
81.7%
14PaLM 540B · proprietary
81.4%
15DeepSeek v2 · 235.7B
80.0%
16Falcon 40B · 41.8B
79.9%
17Llama 2 13B HF · 13.0B
79.6%
18Claude Instant 1.1 · proprietary
78.9%
19Claude Instant 1.2 · proprietary
78.7%
20Llama 13B · proprietary
77.9%
21GLaM (MoE) · proprietary
75.8%
22Mistral 7B v0.1 · 7B
75.2%
23PaLM 2 S · proprietary
75.2%
24Phi 3 Medium 128K Instruct · proprietary
73.9%
25Llama 2 7B HF · 6.7B
73.7%
26Mpt 30B · proprietary
73.6%
27Gemma 7B · 8.5B
72.3%
28Qwen2.5 72B · 72.7B
71.9%
29Text Davinci 001 · proprietary
71.2%
30Llama 7B · 6.7B
71.0%
31Meta Llama 3 8B Instruct · 8.0B
67.7%
32Chinchilla (70B) · proprietary
64.6%
33Falcon 7B · 7.2B
64.6%
34Phi 3 Mini 4k Instruct · 3.8B
64.0%
35Mpt 7B · proprietary
61.6%
36Phi 3 Small 8k Instruct · 7.4B
58.1%
37Gopher (280B) · proprietary
57.2%
38Gemma 2B · 2.5B
53.2%
39Phi 2 · 2.8B
45.2%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

10B100Bmodel size (log scale) →87.6%45.2%DeepSeek v3 · 685B · 82.9%Llama 3.1 405B · 406B · 82.7%DeepSeek v2 · 236B · 80.0%Gemma 7B · 9B · 72.3%Qwen2.5 72B · 73B · 71.9%Llama 7B · 7B · 71.0%Meta Llama 3 8B Instruct · 8B · 67.7%Falcon 7B · 7B · 64.6%Phi 3 Small 8k Instruct · 7B · 58.1%Phi 2 · 3B · 45.2%Gemma 2B · 3B · 53.2%Gemma 2BPhi 3 Mini 4k Instruct · 4B · 64.0%Phi 3 Mini 4k InstructLlama 2 7B HF · 7B · 73.7%Llama 2 7B HFMistral 7B v0.1 · 7B · 75.2%Mistral 7B v0.1Llama 2 13B HF · 13B · 79.6%Llama 2 13B HFFalcon 40B · 42B · 79.9%Falcon 40BMixtral 8x7B v0.1 · 47B · 82.2%Mixtral 8x7B v0.1Llama 2 70B HF · 69B · 87.6%Llama 2 70B HF
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Gemma 2B, 3B, score 53.2% — on the efficiency frontier (best score at its size or smaller).
  • Phi 3 Mini 4k Instruct, 4B, score 64.0% — on the efficiency frontier (best score at its size or smaller).
  • Llama 2 7B HF, 7B, score 73.7% — on the efficiency frontier (best score at its size or smaller).
  • Mistral 7B v0.1, 7B, score 75.2% — on the efficiency frontier (best score at its size or smaller).
  • Llama 2 13B HF, 13B, score 79.6% — on the efficiency frontier (best score at its size or smaller).
  • Falcon 40B, 42B, score 79.9% — on the efficiency frontier (best score at its size or smaller).
  • Mixtral 8x7B v0.1, 47B, score 82.2% — on the efficiency frontier (best score at its size or smaller).
  • Llama 2 70B HF, 69B, score 87.6% — on the efficiency frontier (best score at its size or smaller).

TriviaQA: frequently asked questions

What is the best open LLM on TriviaQA?
Llama 2 70B HF is the top open model on TriviaQA, scoring 87.6%. Among all models tested — including proprietary ones — it ranks #1. That puts it ahead of every proprietary model we track, including Claude 2.0 (Anthropic) at 87.5%.
What's the best TriviaQA model you can run on a 24 GB GPU?
Falcon 40B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 23 GB), scoring 79.9% on TriviaQA.
What's the best TriviaQA model you can run on a 12 GB GPU?
Llama 2 13B HF is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 7 GB), scoring 79.6% on TriviaQA.
Can open models match proprietary models on TriviaQA?
Yes — the best open model (Llama 2 70B HF, 87.6%) matches or beats every proprietary model we track on TriviaQA.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.