Knowledge
OpenBookQA Leaderboard
OpenBookQA is a science-question benchmark modeled on open-book exams: the model is given a small set of core science facts and must combine them with broader common knowledge to answer multi-step questions. It targets reasoning over recall.
Source: epoch21 open models ranked+21 proprietaryData through Apr 2024
Open models ranked on OpenBookQA
# shows rank among open models / rank overall (including proprietary).
| # | Model | Score |
|---|---|---|
| 1 / 1 | Phi 3 Mini 4k Instruct · 3.8B | 88.0% |
| 2 / 2 | Phi 3 Small 8k Instruct · 7.4B | 88.0% |
| 3 / 5 | Mixtral 8x7B v0.1 · 46.7B | 85.8% |
| 4 / 6 | Meta Llama 3 8B Instruct · 8.0B | 82.6% |
| 5 / 7 | Mistral 7B v0.1 · 7B | 79.8% |
| 6 / 8 | Gemma 7B · 8.5B | 78.6% |
| 7 / 9 | Phi 2 · 2.8B | 73.6% |
| 8 / 12 | Falcon 180B · 180B | 64.2% |
| 9 / 14 | Llama 2 70B HF · 69.0B | 60.2% |
| 10 / 16 | Llama 2 7B HF · 6.7B | 58.6% |
| 11 / 20 | Llama 7B · 6.7B | 57.2% |
| 12 / 21 | Llama 2 13B HF · 13.0B | 57.0% |
| 13 / 22 | Falcon 40B · 41.8B | 56.6% |
| 14 / 26 | Falcon 7B · 7.2B | 51.6% |
| 15 / 29 | Xgen 7B 8k Base · 7B | 40.2% |
| 16 / 34 | GPT Neox 20B · 20.7B | 38.8% |
| 17 / 35 | GPT J 6B · 6B | 38.2% |
| 18 / 36 | Phi 1 5 · 1.4B | 37.2% |
| 19 / 37 | Cerebras GPT 13B · 13B | 35.8% |
| 20 / 39 | Stablelm Tuned Alpha 7B · 7B | 32.4% |
| 21 / 42 | Gpt2 Xl · 1.6B | 22.4% |
Score vs model size
Which models give the most quality for their size — the ones worth running locally.
- Phi 1 5, 1B, score 37.2% — on the efficiency frontier (best score at its size or smaller).
- Phi 2, 3B, score 73.6% — on the efficiency frontier (best score at its size or smaller).
- Phi 3 Mini 4k Instruct, 4B, score 88.0% — on the efficiency frontier (best score at its size or smaller).
OpenBookQA: frequently asked questions
- What is the best open LLM on OpenBookQA?
- Phi 3 Mini 4k Instruct is the top open model on OpenBookQA, scoring 88.0%. Among all models tested — including proprietary ones — it ranks #1. That puts it ahead of every proprietary model we track, including Phi 3 Medium 128K Instruct (Microsoft) at 87.4%.
- What's the best OpenBookQA model you can run on a 24 GB GPU?
- Phi 3 Mini 4k Instruct is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 2 GB), scoring 88.0% on OpenBookQA.
- What's the best OpenBookQA model you can run on a 12 GB GPU?
- Phi 3 Mini 4k Instruct is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 2 GB), scoring 88.0% on OpenBookQA.
- Can open models match proprietary models on OpenBookQA?
- Yes — the best open model (Phi 3 Mini 4k Instruct, 88.0%) matches or beats every proprietary model we track on OpenBookQA.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.