Knowledge

OpenBookQA Leaderboard

OpenBookQA is a science-question benchmark modeled on open-book exams: the model is given a small set of core science facts and must combine them with broader common knowledge to answer multi-step questions. It targets reasoning over recall.

Source: epoch21 open models ranked+21 proprietaryData through Apr 2024

All models ranked on OpenBookQA

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1Phi 3 Mini 4k Instruct · 3.8B
88.0%
2Phi 3 Small 8k Instruct · 7.4B
88.0%
3Phi 3 Medium 128K Instruct · proprietary
87.4%
4GPT 3.5 Turbo (Nov 06) · proprietary
86.0%
5Mixtral 8x7B v0.1 · 46.7B
85.8%
6Meta Llama 3 8B Instruct · 8.0B
82.6%
7Mistral 7B v0.1 · 7B
79.8%
8Gemma 7B · 8.5B
78.6%
9Phi 2 · 2.8B
73.6%
10PaLM 540B · proprietary
68.0%
11Text Davinci 001 · proprietary
65.4%
12Falcon 180B · 180B
64.2%
13GLaM (MoE) · proprietary
63.0%
14Llama 2 70B HF · 69.0B
60.2%
15Llama 65B · proprietary
60.2%
16Llama 2 7B HF · 6.7B
58.6%
17Llama 33B · proprietary
58.6%
18Llama 2 34B · proprietary
58.2%
19PaLM 2 M · proprietary
57.4%
20Llama 7B · 6.7B
57.2%
21Llama 2 13B HF · 13.0B
57.0%
22Falcon 40B · 41.8B
56.6%
23Llama 13B · proprietary
56.4%
24PaLM 2 S · proprietary
56.2%
25Mpt 30B · proprietary
52.0%
26Falcon 7B · 7.2B
51.6%
27Mpt 7B · proprietary
51.4%
28PaLM 62B · proprietary
50.4%
29Xgen 7B 8k Base · 7B
40.2%
30RedPajama INCITE 7B Base · proprietary
40.0%
31Opt 13B · proprietary
39.8%
32Dolly v2 12B · proprietary
39.2%
33Open Llama 7B · proprietary
39.0%
34GPT Neox 20B · 20.7B
38.8%
35GPT J 6B · 6B
38.2%
36Phi 1 5 · 1.4B
37.2%
37Cerebras GPT 13B · 13B
35.8%
38Vicuna 13B v1.1 · proprietary
33.0%
39Stablelm Tuned Alpha 7B · 7B
32.4%
40Opt 1.3b · proprietary
24.0%
41GPT Neo 2.7B · proprietary
23.2%
42Gpt2 Xl · 1.6B
22.4%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

10B100Bmodel size (log scale) →88.0%22.4%Phi 3 Small 8k Instruct · 7B · 88.0%Mixtral 8x7B v0.1 · 47B · 85.8%Meta Llama 3 8B Instruct · 8B · 82.6%Mistral 7B v0.1 · 7B · 79.8%Gemma 7B · 9B · 78.6%Falcon 180B · 180B · 64.2%Llama 2 70B HF · 69B · 60.2%Llama 2 7B HF · 7B · 58.6%Llama 7B · 7B · 57.2%Llama 2 13B HF · 13B · 57.0%Falcon 40B · 42B · 56.6%Falcon 7B · 7B · 51.6%Xgen 7B 8k Base · 7B · 40.2%GPT Neox 20B · 21B · 38.8%GPT J 6B · 6B · 38.2%Cerebras GPT 13B · 13B · 35.8%Stablelm Tuned Alpha 7B · 7B · 32.4%Gpt2 Xl · 2B · 22.4%Phi 1 5 · 1B · 37.2%Phi 1 5Phi 2 · 3B · 73.6%Phi 2Phi 3 Mini 4k Instruct · 4B · 88.0%Phi 3 Mini 4k Instruct
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Phi 1 5, 1B, score 37.2% — on the efficiency frontier (best score at its size or smaller).
  • Phi 2, 3B, score 73.6% — on the efficiency frontier (best score at its size or smaller).
  • Phi 3 Mini 4k Instruct, 4B, score 88.0% — on the efficiency frontier (best score at its size or smaller).

OpenBookQA: frequently asked questions

What is the best open LLM on OpenBookQA?
Phi 3 Mini 4k Instruct is the top open model on OpenBookQA, scoring 88.0%. Among all models tested — including proprietary ones — it ranks #1. That puts it ahead of every proprietary model we track, including Phi 3 Medium 128K Instruct (Microsoft) at 87.4%.
What's the best OpenBookQA model you can run on a 24 GB GPU?
Phi 3 Mini 4k Instruct is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 2 GB), scoring 88.0% on OpenBookQA.
What's the best OpenBookQA model you can run on a 12 GB GPU?
Phi 3 Mini 4k Instruct is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 2 GB), scoring 88.0% on OpenBookQA.
Can open models match proprietary models on OpenBookQA?
Yes — the best open model (Phi 3 Mini 4k Instruct, 88.0%) matches or beats every proprietary model we track on OpenBookQA.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.