Knowledge

BoolQ Leaderboard

BoolQ is a reading-comprehension benchmark of naturally occurring yes/no questions, each paired with a short passage that contains the answer. It tests whether a model can extract and reason over information from real text rather than just recall facts.

Source: epoch33 open models ranked+44 proprietaryData through Aug 2024

Open models ranked on BoolQ

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 5StableBeluga2 · 70B
89.4%
2 / 6Falcon 180B · 180B
89.0%
3 / 9Llama 2 70B HF · 69.0B
88.6%
4 / 14Internlm 20B · 20B
87.5%
5 / 15Mistral 7B v0.1 · 7B
87.4%
6 / 18Qwen 14B · 14.2B
86.2%
7 / 21Gemma 2 9B · 9.2B
85.7%
8 / 26Phi 3.5 MoE Instruct · 41.9B
84.6%
9 / 30Gemma 7B · 8.5B
83.2%
10 / 31Mistral 7B Instruct v0.2 · 7.2B
83.2%
11 / 32Falcon 40B · 41.8B
83.1%
12 / 33Llama 3.1 8B Instruct · 8.0B
82.8%
13 / 35Llama 2 13B HF · 13.0B
82.4%
14 / 37Vicuna 13B V1.3 · 13B
80.8%
15 / 40Chatglm2 6B · 6B
79.0%
16 / 43Phi 3.5 Mini Instruct · 3.8B
78.0%
17 / 44Llama 2 7B HF · 6.7B
77.9%
18 / 46Llama 7B · 6.7B
76.5%
19 / 47Qwen 7B · 7.7B
76.4%
20 / 50Phi 1 5 · 1.4B
75.8%
21 / 51Falcon 7B · 7.2B
75.3%
22 / 53Xgen 7B 8k Base · 7B
74.3%
23 / 56Bloom · 176.2B
70.4%
24 / 57Gemma 2B · 2.5B
69.4%
25 / 59Qwen 1 8B · 1.8B
68.0%
26 / 60Baichuan2 13B Base · 13B
67.0%
27 / 62GPT J 6B · 6B
65.4%
28 / 64GPT Neox 20B · 20.7B
64.9%
29 / 65Internlm 7B · 7B
64.1%
30 / 66Baichuan2 7B Base · 7B
63.2%
31 / 69Gpt2 Xl · 1.6B
61.8%
32 / 70Cerebras GPT 13B · 13B
61.1%
33 / 72Stablelm Tuned Alpha 7B · 7B
59.0%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

10B100Bmodel size (log scale) →89.4%59.0%Falcon 180B · 180B · 89.0%Qwen 14B · 14B · 86.2%Gemma 2 9B · 9B · 85.7%Phi 3.5 MoE Instruct · 42B · 84.6%Gemma 7B · 9B · 83.2%Mistral 7B Instruct v0.2 · 7B · 83.2%Falcon 40B · 42B · 83.1%Llama 3.1 8B Instruct · 8B · 82.8%Llama 2 13B HF · 13B · 82.4%Vicuna 13B V1.3 · 13B · 80.8%Llama 2 7B HF · 7B · 77.9%Llama 7B · 7B · 76.5%Qwen 7B · 8B · 76.4%Falcon 7B · 7B · 75.3%Xgen 7B 8k Base · 7B · 74.3%Bloom · 176B · 70.4%Gemma 2B · 3B · 69.4%Qwen 1 8B · 2B · 68.0%Baichuan2 13B Base · 13B · 67.0%GPT J 6B · 6B · 65.4%GPT Neox 20B · 21B · 64.9%Internlm 7B · 7B · 64.1%Baichuan2 7B Base · 7B · 63.2%Gpt2 Xl · 2B · 61.8%Cerebras GPT 13B · 13B · 61.1%Stablelm Tuned Alpha 7B · 7B · 59.0%Phi 1 5 · 1B · 75.8%Phi 1 5Phi 3.5 Mini Instruct · 4B · 78.0%Phi 3.5 Mini InstructChatglm2 6B · 6B · 79.0%Chatglm2 6BMistral 7B v0.1 · 7B · 87.4%Mistral 7B v0.1Internlm 20B · 20B · 87.5%Internlm 20BLlama 2 70B HF · 69B · 88.6%Llama 2 70B HFStableBeluga2 · 70B · 89.4%StableBeluga2
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Phi 1 5, 1B, score 75.8% — on the efficiency frontier (best score at its size or smaller).
  • Phi 3.5 Mini Instruct, 4B, score 78.0% — on the efficiency frontier (best score at its size or smaller).
  • Chatglm2 6B, 6B, score 79.0% — on the efficiency frontier (best score at its size or smaller).
  • Mistral 7B v0.1, 7B, score 87.4% — on the efficiency frontier (best score at its size or smaller).
  • Internlm 20B, 20B, score 87.5% — on the efficiency frontier (best score at its size or smaller).
  • Llama 2 70B HF, 69B, score 88.6% — on the efficiency frontier (best score at its size or smaller).
  • StableBeluga2, 70B, score 89.4% — on the efficiency frontier (best score at its size or smaller).

BoolQ: frequently asked questions

What is the best open LLM on BoolQ?
StableBeluga2 is the top open model on BoolQ, scoring 89.4%. Among all models tested — including proprietary ones — it ranks #5. The top model overall is T5 11B (Google) at 91.2%.
What's the best BoolQ model you can run on a 24 GB GPU?
Internlm 20B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 11 GB), scoring 87.5% on BoolQ.
What's the best BoolQ model you can run on a 12 GB GPU?
Internlm 20B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 11 GB), scoring 87.5% on BoolQ.
Can open models match proprietary models on BoolQ?
Not quite on BoolQ: the strongest proprietary model (T5 11B) scores 91.2%, ahead of the best open model (StableBeluga2) at 89.4% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.