Reasoning
PIQA Leaderboard
PIQA (Physical Interaction QA) asks a model to choose the more sensible of two ways to accomplish an everyday physical task, testing practical, physical-world commonsense that isn't well captured by purely textual reasoning tests.
Source: epoch35 open models ranked+25 proprietaryData through Dec 2024
All models ranked on PIQA
Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.
| # | Model | Score |
|---|---|---|
| 1 | GPT 4o Mini (Jul 18, 2024) · proprietary | 88.7% |
| 2 | Phi 3.5 MoE Instruct · 41.9B | 88.6% |
| 3 | Gemini 1.5 Flash 002 · proprietary | 87.5% |
| 4 | Llama 3.1 405B · 405.9B | 85.9% |
| 5 | Falcon 180B · 180B | 84.9% |
| 6 | DeepSeek V3 Base · proprietary | 84.7% |
| 7 | Inflection 1 · proprietary | 84.2% |
| 8 | DeepSeek v2 · 235.7B | 83.9% |
| 9 | Gemma 2 9B · 9.2B | 83.7% |
| 10 | Mixtral 8x7B v0.1 · 46.7B | 83.6% |
| 11 | Mistral Nemo Base 2407 · proprietary | 83.5% |
| 12 | StableBeluga2 · 70B | 83.3% |
| 13 | Megatron Turing NLG 530B · proprietary | 83.2% |
| 14 | Falcon 40B · 41.8B | 83.0% |
| 15 | Mistral 7B v0.1 · 7B | 83.0% |
| 16 | Llama 2 70B HF · 69.0B | 82.8% |
| 17 | Llama 65B · proprietary | 82.8% |
| 18 | Qwen2.5 72B · 72.7B | 82.6% |
| 19 | Nemotron 4 15B · proprietary | 82.4% |
| 20 | Llama 33B · proprietary | 82.3% |
| 21 | PaLM 540B · proprietary | 82.3% |
| 22 | Text Davinci 001 · proprietary | 82.3% |
| 23 | Mistral 7B Instruct v0.1 · 7.2B | 82.2% |
| 24 | Mistral 7B Instruct v0.2 · 7.2B | 82.2% |
| 25 | Llama 2 34B · proprietary | 81.9% |
| 26 | Mpt 30B · proprietary | 81.9% |
| 27 | Chinchilla (70B) · proprietary | 81.8% |
| 28 | Gopher (280B) · proprietary | 81.8% |
| 29 | Gemma 7B · 8.5B | 81.2% |
| 30 | Llama 3.1 8B Instruct · 8.0B | 81.2% |
| 31 | Phi 3.5 Mini Instruct · 3.8B | 81.0% |
| 32 | Llama 2 13B HF · 13.0B | 80.8% |
| 33 | Mpt 7B · proprietary | 80.6% |
| 34 | PaLM 62B · proprietary | 80.5% |
| 35 | Falcon 7B · 7.2B | 80.3% |
| 36 | Internlm 20B · 20B | 80.3% |
| 37 | Llama 13B · proprietary | 80.1% |
| 38 | Qwen 14B · 14.2B | 79.9% |
| 39 | Llama 7B · 6.7B | 79.8% |
| 40 | Llama 2 7B HF · 6.7B | 78.8% |
| 41 | Baichuan2 13B Base · 13B | 78.1% |
| 42 | Internlm 7B · 7B | 77.9% |
| 43 | Qwen 7B · 7.7B | 77.9% |
| 44 | Vicuna 13B v1.1 · proprietary | 77.4% |
| 45 | Gemma 2B · 2.5B | 77.3% |
| 46 | RedPajama INCITE 7B Base · proprietary | 76.9% |
| 47 | GPT Neox 20B · 20.7B | 76.7% |
| 48 | Baichuan 7B · 7B | 76.2% |
| 49 | Open Llama 7B · proprietary | 76.0% |
| 50 | Opt 13B · proprietary | 75.7% |
| 51 | Xgen 7B 8k Base · 7B | 75.5% |
| 52 | Dolly v2 12B · proprietary | 75.4% |
| 53 | GPT J 6B · 6B | 75.4% |
| 54 | Cerebras GPT 13B · 13B | 73.5% |
| 55 | Qwen 1 8B · 1.8B | 73.3% |
| 56 | GPT Neo 2.7B · proprietary | 72.9% |
| 57 | Gpt2 Xl · 1.6B | 70.5% |
| 58 | Chatglm2 6B · 6B | 69.6% |
| 59 | Opt 1.3b · proprietary | 69.0% |
| 60 | Stablelm Tuned Alpha 7B · 7B | 65.8% |
Score vs model size
Which models give the most quality for their size — the ones worth running locally.
- Gpt2 Xl, 2B, score 70.5% — on the efficiency frontier (best score at its size or smaller).
- Qwen 1 8B, 2B, score 73.3% — on the efficiency frontier (best score at its size or smaller).
- Gemma 2B, 3B, score 77.3% — on the efficiency frontier (best score at its size or smaller).
- Phi 3.5 Mini Instruct, 4B, score 81.0% — on the efficiency frontier (best score at its size or smaller).
- Mistral 7B v0.1, 7B, score 83.0% — on the efficiency frontier (best score at its size or smaller).
- Gemma 2 9B, 9B, score 83.7% — on the efficiency frontier (best score at its size or smaller).
- Phi 3.5 MoE Instruct, 42B, score 88.6% — on the efficiency frontier (best score at its size or smaller).
PIQA: frequently asked questions
- What is the best open LLM on PIQA?
- Phi 3.5 MoE Instruct is the top open model on PIQA, scoring 88.6%. Among all models tested — including proprietary ones — it ranks #2. The top model overall is GPT 4o Mini (Jul 18, 2024) (OpenAI) at 88.7%.
- What's the best PIQA model you can run on a 24 GB GPU?
- Phi 3.5 MoE Instruct is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 23 GB), scoring 88.6% on PIQA.
- What's the best PIQA model you can run on a 12 GB GPU?
- Gemma 2 9B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 5 GB), scoring 83.7% on PIQA.
- Can open models match proprietary models on PIQA?
- Not quite on PIQA: the strongest proprietary model (GPT 4o Mini (Jul 18, 2024)) scores 88.7%, ahead of the best open model (Phi 3.5 MoE Instruct) at 88.6% — but you can run the open one yourself.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.