Knowledge

OpenBookQA Leaderboard

OpenBookQA is a science-question benchmark modeled on open-book exams: the model is given a small set of core science facts and must combine them with broader common knowledge to answer multi-step questions. It targets reasoning over recall.

Source: epoch21 open models ranked+21 proprietaryData through Apr 2024

Open models ranked on OpenBookQA

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 1Phi 3 Mini 4k Instruct · 3.8B
88.0%
2 / 2Phi 3 Small 8k Instruct · 7.4B
88.0%
3 / 5Mixtral 8x7B v0.1 · 46.7B
85.8%
4 / 6Meta Llama 3 8B Instruct · 8.0B
82.6%
5 / 7Mistral 7B v0.1 · 7B
79.8%
6 / 8Gemma 7B · 8.5B
78.6%
7 / 9Phi 2 · 2.8B
73.6%
8 / 12Falcon 180B · 180B
64.2%
9 / 14Llama 2 70B HF · 69.0B
60.2%
10 / 16Llama 2 7B HF · 6.7B
58.6%
11 / 20Llama 7B · 6.7B
57.2%
12 / 21Llama 2 13B HF · 13.0B
57.0%
13 / 22Falcon 40B · 41.8B
56.6%
14 / 26Falcon 7B · 7.2B
51.6%
15 / 29Xgen 7B 8k Base · 7B
40.2%
16 / 34GPT Neox 20B · 20.7B
38.8%
17 / 35GPT J 6B · 6B
38.2%
18 / 36Phi 1 5 · 1.4B
37.2%
19 / 37Cerebras GPT 13B · 13B
35.8%
20 / 39Stablelm Tuned Alpha 7B · 7B
32.4%
21 / 42Gpt2 Xl · 1.6B
22.4%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

10B100Bmodel size (log scale) →88.0%22.4%Phi 3 Small 8k Instruct · 7B · 88.0%Mixtral 8x7B v0.1 · 47B · 85.8%Meta Llama 3 8B Instruct · 8B · 82.6%Mistral 7B v0.1 · 7B · 79.8%Gemma 7B · 9B · 78.6%Falcon 180B · 180B · 64.2%Llama 2 70B HF · 69B · 60.2%Llama 2 7B HF · 7B · 58.6%Llama 7B · 7B · 57.2%Llama 2 13B HF · 13B · 57.0%Falcon 40B · 42B · 56.6%Falcon 7B · 7B · 51.6%Xgen 7B 8k Base · 7B · 40.2%GPT Neox 20B · 21B · 38.8%GPT J 6B · 6B · 38.2%Cerebras GPT 13B · 13B · 35.8%Stablelm Tuned Alpha 7B · 7B · 32.4%Gpt2 Xl · 2B · 22.4%Phi 1 5 · 1B · 37.2%Phi 1 5Phi 2 · 3B · 73.6%Phi 2Phi 3 Mini 4k Instruct · 4B · 88.0%Phi 3 Mini 4k Instruct
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Phi 1 5, 1B, score 37.2% — on the efficiency frontier (best score at its size or smaller).
  • Phi 2, 3B, score 73.6% — on the efficiency frontier (best score at its size or smaller).
  • Phi 3 Mini 4k Instruct, 4B, score 88.0% — on the efficiency frontier (best score at its size or smaller).

OpenBookQA: frequently asked questions

What is the best open LLM on OpenBookQA?
Phi 3 Mini 4k Instruct is the top open model on OpenBookQA, scoring 88.0%. Among all models tested — including proprietary ones — it ranks #1. That puts it ahead of every proprietary model we track, including Phi 3 Medium 128K Instruct (Microsoft) at 87.4%.
What's the best OpenBookQA model you can run on a 24 GB GPU?
Phi 3 Mini 4k Instruct is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 2 GB), scoring 88.0% on OpenBookQA.
What's the best OpenBookQA model you can run on a 12 GB GPU?
Phi 3 Mini 4k Instruct is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 2 GB), scoring 88.0% on OpenBookQA.
Can open models match proprietary models on OpenBookQA?
Yes — the best open model (Phi 3 Mini 4k Instruct, 88.0%) matches or beats every proprietary model we track on OpenBookQA.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.