Knowledge
BoolQ Leaderboard
BoolQ is a reading-comprehension benchmark of naturally occurring yes/no questions, each paired with a short passage that contains the answer. It tests whether a model can extract and reason over information from real text rather than just recall facts.
Source: epoch33 open models ranked+44 proprietaryData through Aug 2024
All models ranked on BoolQ
Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.
| # | Model | Score |
|---|---|---|
| 1 | T5 11B · proprietary | 91.2% |
| 2 | PaLM 2 L · proprietary | 90.9% |
| 3 | T5 3B · proprietary | 89.9% |
| 4 | Inflection 1 · proprietary | 89.7% |
| 5 | StableBeluga2 · 70B | 89.4% |
| 6 | Falcon 180B · 180B | 89.0% |
| 7 | GPT 4o Mini (Jul 18, 2024) · proprietary | 88.7% |
| 8 | PaLM 540B · proprietary | 88.7% |
| 9 | Llama 2 70B HF · 69.0B | 88.6% |
| 10 | PaLM 2 M · proprietary | 88.6% |
| 11 | PaLM 2 S · proprietary | 88.1% |
| 12 | Text Davinci 003 · proprietary | 88.1% |
| 13 | Text Davinci 002 · proprietary | 87.7% |
| 14 | Internlm 20B · 20B | 87.5% |
| 15 | Mistral 7B v0.1 · 7B | 87.4% |
| 16 | Llama 65B · proprietary | 87.1% |
| 17 | GPT 3.5 Turbo (Jun 13) · proprietary | 87.0% |
| 18 | Qwen 14B · 14.2B | 86.2% |
| 19 | Llama 33B · proprietary | 86.1% |
| 20 | Gemini 1.5 Flash 001 · proprietary | 85.8% |
| 21 | Gemma 2 9B · 9.2B | 85.7% |
| 22 | T5 Large · proprietary | 85.4% |
| 23 | Mpt 30B Instruct · proprietary | 85.0% |
| 24 | Megatron Turing NLG 530B · proprietary | 84.8% |
| 25 | PaLM 62B · proprietary | 84.8% |
| 26 | Phi 3.5 MoE Instruct · 41.9B | 84.6% |
| 27 | Chinchilla (70B) · proprietary | 83.7% |
| 28 | Llama 2 34B · proprietary | 83.7% |
| 29 | Vicuna 13B v1.1 · proprietary | 83.5% |
| 30 | Gemma 7B · 8.5B | 83.2% |
| 31 | Mistral 7B Instruct v0.2 · 7.2B | 83.2% |
| 32 | Falcon 40B · 41.8B | 83.1% |
| 33 | Llama 3.1 8B Instruct · 8.0B | 82.8% |
| 34 | Mistral Nemo Base 2407 · proprietary | 82.5% |
| 35 | Llama 2 13B HF · 13.0B | 82.4% |
| 36 | T5 Base · proprietary | 81.4% |
| 37 | Vicuna 13B V1.3 · 13B | 80.8% |
| 38 | Gopher (280B) · proprietary | 79.4% |
| 39 | Opt 175B · proprietary | 79.3% |
| 40 | Chatglm2 6B · 6B | 79.0% |
| 41 | Mpt 30B · proprietary | 79.0% |
| 42 | Llama 13B · proprietary | 78.7% |
| 43 | Phi 3.5 Mini Instruct · 3.8B | 78.0% |
| 44 | Llama 2 7B HF · 6.7B | 77.9% |
| 45 | Text Davinci 001 · proprietary | 77.5% |
| 46 | Llama 7B · 6.7B | 76.5% |
| 47 | Qwen 7B · 7.7B | 76.4% |
| 48 | T5 Small · proprietary | 76.4% |
| 49 | Opt 66B · proprietary | 76.0% |
| 50 | Phi 1 5 · 1.4B | 75.8% |
| 51 | Falcon 7B · 7.2B | 75.3% |
| 52 | Mpt 7B · proprietary | 75.0% |
| 53 | Xgen 7B 8k Base · 7B | 74.3% |
| 54 | Davinci · proprietary | 72.2% |
| 55 | Open Llama 7B · proprietary | 70.6% |
| 56 | Bloom · 176.2B | 70.4% |
| 57 | Gemma 2B · 2.5B | 69.4% |
| 58 | RedPajama INCITE 7B Base · proprietary | 69.3% |
| 59 | Qwen 1 8B · 1.8B | 68.0% |
| 60 | Baichuan2 13B Base · 13B | 67.0% |
| 61 | Curie · proprietary | 65.6% |
| 62 | GPT J 6B · 6B | 65.4% |
| 63 | Opt 13B · proprietary | 65.0% |
| 64 | GPT Neox 20B · 20.7B | 64.9% |
| 65 | Internlm 7B · 7B | 64.1% |
| 66 | Baichuan2 7B Base · 7B | 63.2% |
| 67 | Text Curie 001 · proprietary | 62.0% |
| 68 | GPT Neo 2.7B · proprietary | 61.8% |
| 69 | Gpt2 Xl · 1.6B | 61.8% |
| 70 | Cerebras GPT 13B · 13B | 61.1% |
| 71 | Opt 1.3b · proprietary | 59.6% |
| 72 | Stablelm Tuned Alpha 7B · 7B | 59.0% |
| 73 | Ada · proprietary | 58.1% |
| 74 | Babbage · proprietary | 57.4% |
| 75 | Dolly v2 12B · proprietary | 56.3% |
| 76 | Text Ada 001 · proprietary | 46.4% |
| 77 | Text Babbage 001 · proprietary | 45.1% |
Score vs model size
Which models give the most quality for their size — the ones worth running locally.
- Phi 1 5, 1B, score 75.8% — on the efficiency frontier (best score at its size or smaller).
- Phi 3.5 Mini Instruct, 4B, score 78.0% — on the efficiency frontier (best score at its size or smaller).
- Chatglm2 6B, 6B, score 79.0% — on the efficiency frontier (best score at its size or smaller).
- Mistral 7B v0.1, 7B, score 87.4% — on the efficiency frontier (best score at its size or smaller).
- Internlm 20B, 20B, score 87.5% — on the efficiency frontier (best score at its size or smaller).
- Llama 2 70B HF, 69B, score 88.6% — on the efficiency frontier (best score at its size or smaller).
- StableBeluga2, 70B, score 89.4% — on the efficiency frontier (best score at its size or smaller).
BoolQ: frequently asked questions
- What is the best open LLM on BoolQ?
- StableBeluga2 is the top open model on BoolQ, scoring 89.4%. Among all models tested — including proprietary ones — it ranks #5. The top model overall is T5 11B (Google) at 91.2%.
- What's the best BoolQ model you can run on a 24 GB GPU?
- Internlm 20B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 11 GB), scoring 87.5% on BoolQ.
- What's the best BoolQ model you can run on a 12 GB GPU?
- Internlm 20B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 11 GB), scoring 87.5% on BoolQ.
- Can open models match proprietary models on BoolQ?
- Not quite on BoolQ: the strongest proprietary model (T5 11B) scores 91.2%, ahead of the best open model (StableBeluga2) at 89.4% — but you can run the open one yourself.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.