Reasoning

PIQA Leaderboard

PIQA (Physical Interaction QA) asks a model to choose the more sensible of two ways to accomplish an everyday physical task, testing practical, physical-world commonsense that isn't well captured by purely textual reasoning tests.

Source: epoch35 open models ranked+25 proprietaryData through Dec 2024

Open models ranked on PIQA

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 2Phi 3.5 MoE Instruct · 41.9B
88.6%
2 / 4Llama 3.1 405B · 405.9B
85.9%
3 / 5Falcon 180B · 180B
84.9%
4 / 8DeepSeek v2 · 235.7B
83.9%
5 / 9Gemma 2 9B · 9.2B
83.7%
6 / 10Mixtral 8x7B v0.1 · 46.7B
83.6%
7 / 12StableBeluga2 · 70B
83.3%
8 / 14Falcon 40B · 41.8B
83.0%
9 / 15Mistral 7B v0.1 · 7B
83.0%
10 / 16Llama 2 70B HF · 69.0B
82.8%
11 / 18Qwen2.5 72B · 72.7B
82.6%
12 / 23Mistral 7B Instruct v0.1 · 7.2B
82.2%
13 / 24Mistral 7B Instruct v0.2 · 7.2B
82.2%
14 / 29Gemma 7B · 8.5B
81.2%
15 / 30Llama 3.1 8B Instruct · 8.0B
81.2%
16 / 31Phi 3.5 Mini Instruct · 3.8B
81.0%
17 / 32Llama 2 13B HF · 13.0B
80.8%
18 / 35Falcon 7B · 7.2B
80.3%
19 / 36Internlm 20B · 20B
80.3%
20 / 38Qwen 14B · 14.2B
79.9%
21 / 39Llama 7B · 6.7B
79.8%
22 / 40Llama 2 7B HF · 6.7B
78.8%
23 / 41Baichuan2 13B Base · 13B
78.1%
24 / 42Internlm 7B · 7B
77.9%
25 / 43Qwen 7B · 7.7B
77.9%
26 / 45Gemma 2B · 2.5B
77.3%
27 / 47GPT Neox 20B · 20.7B
76.7%
28 / 48Baichuan 7B · 7B
76.2%
29 / 51Xgen 7B 8k Base · 7B
75.5%
30 / 53GPT J 6B · 6B
75.4%
31 / 54Cerebras GPT 13B · 13B
73.5%
32 / 55Qwen 1 8B · 1.8B
73.3%
33 / 57Gpt2 Xl · 1.6B
70.5%
34 / 58Chatglm2 6B · 6B
69.6%
35 / 60Stablelm Tuned Alpha 7B · 7B
65.8%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

10B100Bmodel size (log scale) →88.6%65.8%Llama 3.1 405B · 406B · 85.9%Falcon 180B · 180B · 84.9%DeepSeek v2 · 236B · 83.9%Mixtral 8x7B v0.1 · 47B · 83.6%StableBeluga2 · 70B · 83.3%Falcon 40B · 42B · 83.0%Llama 2 70B HF · 69B · 82.8%Qwen2.5 72B · 73B · 82.6%Mistral 7B Instruct v0.1 · 7B · 82.2%Mistral 7B Instruct v0.2 · 7B · 82.2%Gemma 7B · 9B · 81.2%Llama 3.1 8B Instruct · 8B · 81.2%Llama 2 13B HF · 13B · 80.8%Internlm 20B · 20B · 80.3%Falcon 7B · 7B · 80.3%Qwen 14B · 14B · 79.9%Llama 7B · 7B · 79.8%Llama 2 7B HF · 7B · 78.8%Baichuan2 13B Base · 13B · 78.1%Internlm 7B · 7B · 77.9%Qwen 7B · 8B · 77.9%GPT Neox 20B · 21B · 76.7%Baichuan 7B · 7B · 76.2%Xgen 7B 8k Base · 7B · 75.5%GPT J 6B · 6B · 75.4%Cerebras GPT 13B · 13B · 73.5%Chatglm2 6B · 6B · 69.6%Stablelm Tuned Alpha 7B · 7B · 65.8%Gpt2 Xl · 2B · 70.5%Gpt2 XlQwen 1 8B · 2B · 73.3%Qwen 1 8BGemma 2B · 3B · 77.3%Gemma 2BPhi 3.5 Mini Instruct · 4B · 81.0%Phi 3.5 Mini InstructMistral 7B v0.1 · 7B · 83.0%Mistral 7B v0.1Gemma 2 9B · 9B · 83.7%Gemma 2 9BPhi 3.5 MoE Instruct · 42B · 88.6%Phi 3.5 MoE Instruct
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Gpt2 Xl, 2B, score 70.5% — on the efficiency frontier (best score at its size or smaller).
  • Qwen 1 8B, 2B, score 73.3% — on the efficiency frontier (best score at its size or smaller).
  • Gemma 2B, 3B, score 77.3% — on the efficiency frontier (best score at its size or smaller).
  • Phi 3.5 Mini Instruct, 4B, score 81.0% — on the efficiency frontier (best score at its size or smaller).
  • Mistral 7B v0.1, 7B, score 83.0% — on the efficiency frontier (best score at its size or smaller).
  • Gemma 2 9B, 9B, score 83.7% — on the efficiency frontier (best score at its size or smaller).
  • Phi 3.5 MoE Instruct, 42B, score 88.6% — on the efficiency frontier (best score at its size or smaller).

PIQA: frequently asked questions

What is the best open LLM on PIQA?
Phi 3.5 MoE Instruct is the top open model on PIQA, scoring 88.6%. Among all models tested — including proprietary ones — it ranks #2. The top model overall is GPT 4o Mini (Jul 18, 2024) (OpenAI) at 88.7%.
What's the best PIQA model you can run on a 24 GB GPU?
Phi 3.5 MoE Instruct is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 23 GB), scoring 88.6% on PIQA.
What's the best PIQA model you can run on a 12 GB GPU?
Gemma 2 9B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 5 GB), scoring 83.7% on PIQA.
Can open models match proprietary models on PIQA?
Not quite on PIQA: the strongest proprietary model (GPT 4o Mini (Jul 18, 2024)) scores 88.7%, ahead of the best open model (Phi 3.5 MoE Instruct) at 88.6% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.