Reasoning

PIQA Leaderboard

PIQA (Physical Interaction QA) asks a model to choose the more sensible of two ways to accomplish an everyday physical task, testing practical, physical-world commonsense that isn't well captured by purely textual reasoning tests.

Source: epoch35 open models ranked+25 proprietaryData through Dec 2024

All models ranked on PIQA

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1GPT 4o Mini (Jul 18, 2024) · proprietary
88.7%
2Phi 3.5 MoE Instruct · 41.9B
88.6%
3Gemini 1.5 Flash 002 · proprietary
87.5%
4Llama 3.1 405B · 405.9B
85.9%
5Falcon 180B · 180B
84.9%
6DeepSeek V3 Base · proprietary
84.7%
7Inflection 1 · proprietary
84.2%
8DeepSeek v2 · 235.7B
83.9%
9Gemma 2 9B · 9.2B
83.7%
10Mixtral 8x7B v0.1 · 46.7B
83.6%
11Mistral Nemo Base 2407 · proprietary
83.5%
12StableBeluga2 · 70B
83.3%
13Megatron Turing NLG 530B · proprietary
83.2%
14Falcon 40B · 41.8B
83.0%
15Mistral 7B v0.1 · 7B
83.0%
16Llama 2 70B HF · 69.0B
82.8%
17Llama 65B · proprietary
82.8%
18Qwen2.5 72B · 72.7B
82.6%
19Nemotron 4 15B · proprietary
82.4%
20Llama 33B · proprietary
82.3%
21PaLM 540B · proprietary
82.3%
22Text Davinci 001 · proprietary
82.3%
23Mistral 7B Instruct v0.1 · 7.2B
82.2%
24Mistral 7B Instruct v0.2 · 7.2B
82.2%
25Llama 2 34B · proprietary
81.9%
26Mpt 30B · proprietary
81.9%
27Chinchilla (70B) · proprietary
81.8%
28Gopher (280B) · proprietary
81.8%
29Gemma 7B · 8.5B
81.2%
30Llama 3.1 8B Instruct · 8.0B
81.2%
31Phi 3.5 Mini Instruct · 3.8B
81.0%
32Llama 2 13B HF · 13.0B
80.8%
33Mpt 7B · proprietary
80.6%
34PaLM 62B · proprietary
80.5%
35Falcon 7B · 7.2B
80.3%
36Internlm 20B · 20B
80.3%
37Llama 13B · proprietary
80.1%
38Qwen 14B · 14.2B
79.9%
39Llama 7B · 6.7B
79.8%
40Llama 2 7B HF · 6.7B
78.8%
41Baichuan2 13B Base · 13B
78.1%
42Internlm 7B · 7B
77.9%
43Qwen 7B · 7.7B
77.9%
44Vicuna 13B v1.1 · proprietary
77.4%
45Gemma 2B · 2.5B
77.3%
46RedPajama INCITE 7B Base · proprietary
76.9%
47GPT Neox 20B · 20.7B
76.7%
48Baichuan 7B · 7B
76.2%
49Open Llama 7B · proprietary
76.0%
50Opt 13B · proprietary
75.7%
51Xgen 7B 8k Base · 7B
75.5%
52Dolly v2 12B · proprietary
75.4%
53GPT J 6B · 6B
75.4%
54Cerebras GPT 13B · 13B
73.5%
55Qwen 1 8B · 1.8B
73.3%
56GPT Neo 2.7B · proprietary
72.9%
57Gpt2 Xl · 1.6B
70.5%
58Chatglm2 6B · 6B
69.6%
59Opt 1.3b · proprietary
69.0%
60Stablelm Tuned Alpha 7B · 7B
65.8%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

10B100Bmodel size (log scale) →88.6%65.8%Llama 3.1 405B · 406B · 85.9%Falcon 180B · 180B · 84.9%DeepSeek v2 · 236B · 83.9%Mixtral 8x7B v0.1 · 47B · 83.6%StableBeluga2 · 70B · 83.3%Falcon 40B · 42B · 83.0%Llama 2 70B HF · 69B · 82.8%Qwen2.5 72B · 73B · 82.6%Mistral 7B Instruct v0.1 · 7B · 82.2%Mistral 7B Instruct v0.2 · 7B · 82.2%Gemma 7B · 9B · 81.2%Llama 3.1 8B Instruct · 8B · 81.2%Llama 2 13B HF · 13B · 80.8%Internlm 20B · 20B · 80.3%Falcon 7B · 7B · 80.3%Qwen 14B · 14B · 79.9%Llama 7B · 7B · 79.8%Llama 2 7B HF · 7B · 78.8%Baichuan2 13B Base · 13B · 78.1%Internlm 7B · 7B · 77.9%Qwen 7B · 8B · 77.9%GPT Neox 20B · 21B · 76.7%Baichuan 7B · 7B · 76.2%Xgen 7B 8k Base · 7B · 75.5%GPT J 6B · 6B · 75.4%Cerebras GPT 13B · 13B · 73.5%Chatglm2 6B · 6B · 69.6%Stablelm Tuned Alpha 7B · 7B · 65.8%Gpt2 Xl · 2B · 70.5%Gpt2 XlQwen 1 8B · 2B · 73.3%Qwen 1 8BGemma 2B · 3B · 77.3%Gemma 2BPhi 3.5 Mini Instruct · 4B · 81.0%Phi 3.5 Mini InstructMistral 7B v0.1 · 7B · 83.0%Mistral 7B v0.1Gemma 2 9B · 9B · 83.7%Gemma 2 9BPhi 3.5 MoE Instruct · 42B · 88.6%Phi 3.5 MoE Instruct
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Gpt2 Xl, 2B, score 70.5% — on the efficiency frontier (best score at its size or smaller).
  • Qwen 1 8B, 2B, score 73.3% — on the efficiency frontier (best score at its size or smaller).
  • Gemma 2B, 3B, score 77.3% — on the efficiency frontier (best score at its size or smaller).
  • Phi 3.5 Mini Instruct, 4B, score 81.0% — on the efficiency frontier (best score at its size or smaller).
  • Mistral 7B v0.1, 7B, score 83.0% — on the efficiency frontier (best score at its size or smaller).
  • Gemma 2 9B, 9B, score 83.7% — on the efficiency frontier (best score at its size or smaller).
  • Phi 3.5 MoE Instruct, 42B, score 88.6% — on the efficiency frontier (best score at its size or smaller).

PIQA: frequently asked questions

What is the best open LLM on PIQA?
Phi 3.5 MoE Instruct is the top open model on PIQA, scoring 88.6%. Among all models tested — including proprietary ones — it ranks #2. The top model overall is GPT 4o Mini (Jul 18, 2024) (OpenAI) at 88.7%.
What's the best PIQA model you can run on a 24 GB GPU?
Phi 3.5 MoE Instruct is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 23 GB), scoring 88.6% on PIQA.
What's the best PIQA model you can run on a 12 GB GPU?
Gemma 2 9B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 5 GB), scoring 83.7% on PIQA.
Can open models match proprietary models on PIQA?
Not quite on PIQA: the strongest proprietary model (GPT 4o Mini (Jul 18, 2024)) scores 88.7%, ahead of the best open model (Phi 3.5 MoE Instruct) at 88.6% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.