Reasoning
LMCA Leaderboard
LMCA (Language Model Conceptual Argumentation) scores models on expert-rated arguments across philosophy, decision theory, and AI risk — domains where there is no empirical feedback loop to check an answer against. It is part of the Conceptual Reasoning Index, built with Anthropic by researchers Emery Cooper and Caspar Oesterheld.
Source: epoch53 open models ranked+119 proprietaryData through Sep 2026
Open models ranked on LMCA
# shows rank among open models / rank overall (including proprietary).
| # | Model | Score |
|---|---|---|
| 1 / 29 | GLM 5.3 · 753.3B | 55.5% |
| 2 / 38 | Kimi K3 · 2779.9B | 52.7% |
| 3 / 63 | DeepSeek V4.1 Flash · 763.2B | 47.0% |
| 4 / 67 | GLM 5.2 · 753.3B | 45.8% |
| 5 / 68 | DeepSeek V4 Pro 0813 · 1650.5B | 45.5% |
| 6 / 79 | DeepSeek V4 Flash 0731 · 304.2B | 41.7% |
| 7 / 81 | Qwen3.8 27B · 27.8B | 41.4% |
| 8 / 82 | DeepSeek V4 Pro · 1598.8B | 41.2% |
| 9 / 86 | Gemma 4 31B IT · 31.3B | 39.3% |
| 10 / 91 | Qwen3.5 397B A17B · 403.4B | 37.9% |
| 11 / 93 | Inkling · 952.4B | 37.6% |
| 12 / 96 | Kimi K2.6 · 1026.9B | 37.3% |
| 13 / 98 | NVIDIA Nemotron 3 Ultra 550B A55B BF16 · 560.5B | 36.9% |
| 14 / 101 | DeepSeek V4 Flash · 290.9B | 35.9% |
| 15 / 104 | Qwen3.6 27B · 27.8B | 34.5% |
| 16 / 106 | Qwen3.5 27B · 27.8B | 34.0% |
| 17 / 107 | MiniMax M3 · 427.0B | 33.7% |
| 18 / 109 | Qwen3.5 122B A10B · 125.1B | 32.2% |
| 19 / 112 | Qwen3.6 35B A3B · 36.0B | 29.7% |
| 20 / 113 | Gemma 4 26B A4B IT · 25.8B | 29.7% |
| 21 / 114 | MiMo V2.5 Pro · 1023.2B | 29.5% |
| 22 / 115 | Qwen3.5 35B A3B · 36.0B | 29.5% |
| 23 / 116 | Qwen3 235B A22B Thinking 2507 · 235.1B | 29.3% |
| 24 / 118 | DeepSeek V3.2 Exp · 685.4B | 29.1% |
| 25 / 120 | DeepSeek V3.2 · 685.4B | 28.8% |
| 26 / 121 | DeepSeek V3.1 Terminus · 684.5B | 28.6% |
| 27 / 127 | Qwen3 235B A22B · 235.1B | 25.0% |
| 28 / 128 | Qwen3.5 9B · 9.7B | 24.5% |
| 29 / 129 | DeepSeek V3.1 · 684.5B | 24.3% |
| 30 / 131 | Qwen3 235B A22B Instruct 2507 · 235.1B | 23.5% |
| 31 / 132 | Qwen3 30B A3B Instruct 2507 · 30.5B | 22.4% |
| 32 / 134 | GPT OSS 120B · 116.8B | 22.1% |
| 33 / 137 | Qwen3 30B A3B Thinking 2507 · 30.5B | 19.5% |
| 34 / 140 | Qwen3 14B · 14.8B | 18.2% |
| 35 / 142 | Llama 3.3 70B Instruct · 70.6B | 17.5% |
| 36 / 143 | Qwen3 32B · 32.8B | 17.3% |
| 37 / 146 | Mistral Large 3 675B Instruct 2512 · 675B | 16.7% |
| 38 / 148 | Llama 4 Maverick 17B 128E Instruct · 401.6B | 15.9% |
| 39 / 149 | Qwen3 30B A3B · 30.5B | 15.8% |
| 40 / 150 | DeepSeek v3 0324 · 684.5B | 15.5% |
| 41 / 152 | Llama 3.1 70B Instruct · 70.6B | 14.8% |
| 42 / 153 | GPT OSS 20B · 20.9B | 14.5% |
| 43 / 154 | Qwen2.5 72B Instruct · 72.7B | 13.4% |
| 44 / 155 | Gemma 3 27B IT · 27.4B | 12.3% |
| 45 / 156 | Llama 4 Scout 17B 16E Instruct · 108.6B | 12.0% |
| 46 / 158 | C4ai Command A 03 2025 · 111.1B | 10.3% |
| 47 / 161 | C4ai Command R 08 2024 · 32.3B | 9.2% |
| 48 / 163 | Qwen3 8B · 8.2B | 8.8% |
| 49 / 166 | Qwen2.5 7B Instruct · 7.6B | 6.4% |
| 50 / 169 | Llama 3.1 8B Instruct · 8.0B | 5.4% |
| 51 / 170 | C4ai Command R Plus 08 2024 · 103.8B | 5.0% |
| 52 / 171 | Gemma 3 12B IT · 12.2B | 4.5% |
| 53 / 172 | Gemma 3 4B IT · 4.3B | 2.8% |
Score vs model size
Which models give the most quality for their size — the ones worth running locally.
- Gemma 3 4B IT, 4B, score 2.8% — on the efficiency frontier (best score at its size or smaller).
- Qwen2.5 7B Instruct, 8B, score 6.4% — on the efficiency frontier (best score at its size or smaller).
- Qwen3 8B, 8B, score 8.8% — on the efficiency frontier (best score at its size or smaller).
- Qwen3.5 9B, 10B, score 24.5% — on the efficiency frontier (best score at its size or smaller).
- Gemma 4 26B A4B IT, 26B, score 29.7% — on the efficiency frontier (best score at its size or smaller).
- Qwen3.8 27B, 28B, score 41.4% — on the efficiency frontier (best score at its size or smaller).
- DeepSeek V4 Flash 0731, 304B, score 41.7% — on the efficiency frontier (best score at its size or smaller).
- GLM 5.3, 753B, score 55.5% — on the efficiency frontier (best score at its size or smaller).
LMCA: frequently asked questions
- What is the best open LLM on LMCA?
- GLM 5.3 is the top open model on LMCA, scoring 55.5%. Among all models tested — including proprietary ones — it ranks #29. The top model overall is Claude Fable 5.1 (unspecified) (Anthropic) at 65.5%.
- What's the best LMCA model you can run on a 24 GB GPU?
- Qwen3.8 27B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 15 GB), scoring 41.4% on LMCA.
- What's the best LMCA model you can run on a 12 GB GPU?
- Qwen3.5 9B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 5 GB), scoring 24.5% on LMCA.
- Can open models match proprietary models on LMCA?
- Not quite on LMCA: the strongest proprietary model (Claude Fable 5.1 (unspecified)) scores 65.5%, ahead of the best open model (GLM 5.3) at 55.5% — but you can run the open one yourself.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.