Reasoning
LMCA Leaderboard
LMCA (Language Model Conceptual Argumentation) scores models on expert-rated arguments across philosophy, decision theory, and AI risk — domains where there is no empirical feedback loop to check an answer against. It is part of the Conceptual Reasoning Index, built with Anthropic by researchers Emery Cooper and Caspar Oesterheld.
Source: epoch53 open models ranked+119 proprietaryData through Sep 2026
All models ranked on LMCA
Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.
| # | Model | Score |
|---|---|---|
| 1 | Claude Fable 5.1 (unspecified) · proprietary | 65.5% |
| 2 | Claude Fable 5.1 (medium) · proprietary | 65.3% |
| 3 | Claude Fable 5.1 (high) · proprietary | 65.1% |
| 4 | Claude Fable 5.1 (xhigh) · proprietary | 65.1% |
| 5 | Claude Opus 5 (xhigh) · proprietary | 64.5% |
| 6 | GPT 6 Astra (xhigh) · proprietary | 64.4% |
| 7 | GPT 6 Astra (unspecified) · proprietary | 64.1% |
| 8 | GPT 6 Astra (high) · proprietary | 63.8% |
| 9 | Claude Opus 5 (high) · proprietary | 63.7% |
| 10 | Claude Fable 5.1 (low) · proprietary | 63.5% |
| 11 | Claude Opus 5 Max · proprietary | 63.3% |
| 12 | Claude Opus 5 (medium) · proprietary | 62.9% |
| 13 | GPT 6 Astra (medium) · proprietary | 62.9% |
| 14 | Claude Opus 5 (low) · proprietary | 62.6% |
| 15 | GPT 6 Astra (low) · proprietary | 62.2% |
| 16 | Claude Fable 5 (high) · proprietary | 61.1% |
| 17 | Claude Fable 5 (xhigh) · proprietary | 61.0% |
| 18 | Claude Fable 5 Max · proprietary | 60.3% |
| 19 | Claude Fable 5 (medium) · proprietary | 60.3% |
| 20 | Claude Fable 5 (low) · proprietary | 59.6% |
| 21 | GPT 5.6 Sol Promax · proprietary | 59.2% |
| 22 | GPT 5.6 Sol (xhigh) · proprietary | 58.5% |
| 23 | GPT 5.6 Sol Max · proprietary | 58.4% |
| 24 | Claude Opus 4.8 Max · proprietary | 57.5% |
| 25 | GPT 5.6 Sol (high) · proprietary | 56.9% |
| 26 | GPT 5.6 Sol (medium) · proprietary | 56.4% |
| 27 | GPT 5.6 Sol (low) · proprietary | 55.9% |
| 28 | Claude Opus 4.6 Max · proprietary | 55.8% |
| 29 | GLM 5.3 · 753.3B | 55.5% |
| 30 | GPT 5.6 Terra Max · proprietary | 55.0% |
| 31 | GPT 5.5 (xhigh) · proprietary | 54.3% |
| 32 | GPT 5.5 Pro (xhigh) · proprietary | 53.9% |
| 33 | Muse Spark 1.3 (unspecified) · proprietary | 53.9% |
| 34 | Gemini 3.1 Pro Preview (high) · proprietary | 53.8% |
| 35 | Gemini 3.8 Flash (medium) · proprietary | 52.9% |
| 36 | Gemini 3.1 Pro Preview (medium) · proprietary | 52.8% |
| 37 | GPT 5.6 Sol None · proprietary | 52.8% |
| 38 | Kimi K3 · 2779.9B | 52.7% |
| 39 | GPT 5.6 Terra (xhigh) · proprietary | 52.2% |
| 40 | Claude Opus 4.7 Max · proprietary | 52.2% |
| 41 | GPT 5.4 (Mar 05, 2026, xhigh) · proprietary | 52.0% |
| 42 | Gemini 3.8 Flash (unspecified) · proprietary | 51.5% |
| 43 | GPT 5.6 Terra (high) · proprietary | 51.4% |
| 44 | Gemini 3.1 Pro Preview (low) · proprietary | 51.2% |
| 45 | Gemini 3.8 Flash (low) · proprietary | 51.2% |
| 46 | Gemini 3.7 Flash (high) · proprietary | 50.4% |
| 47 | Claude Sonnet 5 (high) · proprietary | 50.0% |
| 48 | Muse Spark 1.1 (high) · proprietary | 49.9% |
| 49 | Claude Sonnet 5 (medium) · proprietary | 49.7% |
| 50 | Claude Sonnet 5 (low) · proprietary | 49.6% |
| 51 | GPT 5.6 Terra (medium) · proprietary | 49.6% |
| 52 | Claude Sonnet 5 Max · proprietary | 49.3% |
| 53 | GPT 5.6 Terra (low) · proprietary | 49.3% |
| 54 | Gemini 3.7 Flash (medium) · proprietary | 49.2% |
| 55 | Claude Sonnet 5 (xhigh) · proprietary | 48.8% |
| 56 | GPT 5.6 Luna (xhigh) · proprietary | 48.5% |
| 57 | GPT 5.6 Luna Max · proprietary | 48.5% |
| 58 | Grok 4.6 (xhigh) · proprietary | 48.5% |
| 59 | Muse Spark 1.2 (xhigh) · proprietary | 48.4% |
| 60 | GPT 5.6 Luna (high) · proprietary | 47.4% |
| 61 | Gemini 3.5 Flash (high) · proprietary | 47.1% |
| 62 | Gemini 3.7 Flash (low) · proprietary | 47.1% |
| 63 | DeepSeek V4.1 Flash · 763.2B | 47.0% |
| 64 | Claude Sonnet 4.6 Max · proprietary | 46.5% |
| 65 | GPT 5.6 Terra None · proprietary | 46.5% |
| 66 | Qwen3.8 Max (xhigh) · proprietary | 46.2% |
| 67 | GLM 5.2 · 753.3B | 45.8% |
| 68 | DeepSeek V4 Pro 0813 · 1650.5B | 45.5% |
| 69 | Grok 4.5 (high) · proprietary | 45.2% |
| 70 | Gemini 3.6 Flash (high) · proprietary | 44.9% |
| 71 | Claude Opus 4.5 (Nov 01, 2025, high) · proprietary | 44.5% |
| 72 | GPT 5.6 Luna (medium) · proprietary | 44.3% |
| 73 | Qwen3.7 Max · proprietary | 44.0% |
| 74 | GPT 5.2 (Dec 11, 2025, xhigh) · proprietary | 43.9% |
| 75 | GPT 5.1 (Nov 13, 2025, high) · proprietary | 43.9% |
| 76 | GPT 5.6 Luna (low) · proprietary | 43.8% |
| 77 | Gemini 3 Flash Preview (high) · proprietary | 43.1% |
| 78 | Qwen3.6 Max Preview · proprietary | 42.5% |
| 79 | DeepSeek V4 Flash 0731 · 304.2B | 41.7% |
| 80 | GPT 5.6 Luna None · proprietary | 41.4% |
| 81 | Qwen3.8 27B · 27.8B | 41.4% |
| 82 | DeepSeek V4 Pro · 1598.8B | 41.2% |
| 83 | GPT 5.4 Mini (Mar 17, 2026, xhigh) · proprietary | 40.8% |
| 84 | GPT 5 (Aug 07, 2025, high) · proprietary | 40.0% |
| 85 | O3 (Apr 16, 2025, high) · proprietary | 39.7% |
| 86 | Gemma 4 31B IT · 31.3B | 39.3% |
| 87 | Claude Sonnet 4.5 (Sep 29, 2025, unspecified) · proprietary | 38.8% |
| 88 | Grok 4.20 · proprietary | 38.7% |
| 89 | O3 Pro (Jun 10, 2025, high) · proprietary | 38.5% |
| 90 | Grok 4.3 (high) · proprietary | 38.3% |
| 91 | Qwen3.5 397B A17B · 403.4B | 37.9% |
| 92 | Gemini 3.5 Flash Lite (high) · proprietary | 37.6% |
| 93 | Inkling · 952.4B | 37.6% |
| 94 | Qwen3.7 Plus · proprietary | 37.6% |
| 95 | Claude Opus 4 (May 14, 2025, unspecified) · proprietary | 37.4% |
| 96 | Kimi K2.6 · 1026.9B | 37.3% |
| 97 | Claude Opus 4.1 (Aug 05, 2025, unspecified) · proprietary | 37.1% |
| 98 | NVIDIA Nemotron 3 Ultra 550B A55B BF16 · 560.5B | 36.9% |
| 99 | GPT 5.4 Nano (Mar 17, 2026, xhigh) · proprietary | 36.9% |
| 100 | Qwen3.5 Plus · proprietary | 36.4% |
| 101 | DeepSeek V4 Flash · 290.9B | 35.9% |
| 102 | Gemini 3.1 Flash Lite (high) · proprietary | 35.0% |
| 103 | Gemini 2.5 Pro · proprietary | 34.8% |
| 104 | Qwen3.6 27B · 27.8B | 34.5% |
| 105 | GPT 5 Mini (Aug 07, 2025, high) · proprietary | 34.2% |
| 106 | Qwen3.5 27B · 27.8B | 34.0% |
| 107 | MiniMax M3 · 427.0B | 33.7% |
| 108 | Qwen3.6 Plus · proprietary | 33.1% |
| 109 | Qwen3.5 122B A10B · 125.1B | 32.2% |
| 110 | Qwen3.6 Flash · proprietary | 31.0% |
| 111 | Claude Haiku 4.5 (Oct 01, 2025, unspecified) · proprietary | 30.9% |
| 112 | Qwen3.6 35B A3B · 36.0B | 29.7% |
| 113 | Gemma 4 26B A4B IT · 25.8B | 29.7% |
| 114 | MiMo V2.5 Pro · 1023.2B | 29.5% |
| 115 | Qwen3.5 35B A3B · 36.0B | 29.5% |
| 116 | Qwen3 235B A22B Thinking 2507 · 235.1B | 29.3% |
| 117 | Qwen3.5 Flash · proprietary | 29.1% |
| 118 | DeepSeek V3.2 Exp · 685.4B | 29.1% |
| 119 | Claude Sonnet 4 (May 14, 2025, unspecified) · proprietary | 29.0% |
| 120 | DeepSeek V3.2 · 685.4B | 28.8% |
| 121 | DeepSeek V3.1 Terminus · 684.5B | 28.6% |
| 122 | Qwen3 Max (Sep 23, 2025) · proprietary | 28.3% |
| 123 | Gemini 2.5 Flash · proprietary | 27.5% |
| 124 | O4 Mini (Apr 16, 2025, high) · proprietary | 26.5% |
| 125 | Mistral Medium 2604 · proprietary | 26.1% |
| 126 | GPT 4.1 (Apr 14, 2025) · proprietary | 25.6% |
| 127 | Qwen3 235B A22B · 235.1B | 25.0% |
| 128 | Qwen3.5 9B · 9.7B | 24.5% |
| 129 | DeepSeek V3.1 · 684.5B | 24.3% |
| 130 | Qwen Plus (Apr 28, 2025) · proprietary | 24.0% |
| 131 | Qwen3 235B A22B Instruct 2507 · 235.1B | 23.5% |
| 132 | Qwen3 30B A3B Instruct 2507 · 30.5B | 22.4% |
| 133 | O1 (Dec 17, 2024, high) · proprietary | 22.3% |
| 134 | GPT OSS 120B · 116.8B | 22.1% |
| 135 | GPT 4.1 Mini (Apr 14, 2025) · proprietary | 21.1% |
| 136 | Mistral Small 2603 · proprietary | 20.6% |
| 137 | Qwen3 30B A3B Thinking 2507 · 30.5B | 19.5% |
| 138 | Mistral Medium 2508 · proprietary | 19.1% |
| 139 | O3 Mini (Jan 31, 2025, high) · proprietary | 19.0% |
| 140 | Qwen3 14B · 14.8B | 18.2% |
| 141 | Gemini 2.5 Flash Lite Preview Thinking (Jun 17) · proprietary | 18.1% |
| 142 | Llama 3.3 70B Instruct · 70.6B | 17.5% |
| 143 | Qwen3 32B · 32.8B | 17.3% |
| 144 | GPT 4 (Jun 13) · proprietary | 17.1% |
| 145 | Claude 3 Opus (Feb 29, 2024) · proprietary | 17.0% |
| 146 | Mistral Large 3 675B Instruct 2512 · 675B | 16.7% |
| 147 | GPT 4o (May 13, 2024) · proprietary | 16.6% |
| 148 | Llama 4 Maverick 17B 128E Instruct · 401.6B | 15.9% |
| 149 | Qwen3 30B A3B · 30.5B | 15.8% |
| 150 | DeepSeek v3 0324 · 684.5B | 15.5% |
| 151 | DeepSeek Chat · proprietary | 15.2% |
| 152 | Llama 3.1 70B Instruct · 70.6B | 14.8% |
| 153 | GPT OSS 20B · 20.9B | 14.5% |
| 154 | Qwen2.5 72B Instruct · 72.7B | 13.4% |
| 155 | Gemma 3 27B IT · 27.4B | 12.3% |
| 156 | Llama 4 Scout 17B 16E Instruct · 108.6B | 12.0% |
| 157 | GPT 4o Mini (Jul 18, 2024) · proprietary | 10.4% |
| 158 | C4ai Command A 03 2025 · 111.1B | 10.3% |
| 159 | GPT 4 Turbo · proprietary | 9.8% |
| 160 | GPT 3.5 Turbo (Jan 25) · proprietary | 9.7% |
| 161 | C4ai Command R 08 2024 · 32.3B | 9.2% |
| 162 | Claude 3 Haiku (Mar 07, 2024) · proprietary | 8.8% |
| 163 | Qwen3 8B · 8.2B | 8.8% |
| 164 | GPT 5 Nano (Aug 07, 2025, high) · proprietary | 7.9% |
| 165 | Gemma 2 27B · proprietary | 7.1% |
| 166 | Qwen2.5 7B Instruct · 7.6B | 6.4% |
| 167 | Ministral 3B 2410 · proprietary | 5.5% |
| 168 | GPT 4.1 Nano (Apr 14, 2025) · proprietary | 5.5% |
| 169 | Llama 3.1 8B Instruct · 8.0B | 5.4% |
| 170 | C4ai Command R Plus 08 2024 · 103.8B | 5.0% |
| 171 | Gemma 3 12B IT · 12.2B | 4.5% |
| 172 | Gemma 3 4B IT · 4.3B | 2.8% |
Score vs model size
Which models give the most quality for their size — the ones worth running locally.
- Gemma 3 4B IT, 4B, score 2.8% — on the efficiency frontier (best score at its size or smaller).
- Qwen2.5 7B Instruct, 8B, score 6.4% — on the efficiency frontier (best score at its size or smaller).
- Qwen3 8B, 8B, score 8.8% — on the efficiency frontier (best score at its size or smaller).
- Qwen3.5 9B, 10B, score 24.5% — on the efficiency frontier (best score at its size or smaller).
- Gemma 4 26B A4B IT, 26B, score 29.7% — on the efficiency frontier (best score at its size or smaller).
- Qwen3.8 27B, 28B, score 41.4% — on the efficiency frontier (best score at its size or smaller).
- DeepSeek V4 Flash 0731, 304B, score 41.7% — on the efficiency frontier (best score at its size or smaller).
- GLM 5.3, 753B, score 55.5% — on the efficiency frontier (best score at its size or smaller).
LMCA: frequently asked questions
- What is the best open LLM on LMCA?
- GLM 5.3 is the top open model on LMCA, scoring 55.5%. Among all models tested — including proprietary ones — it ranks #29. The top model overall is Claude Fable 5.1 (unspecified) (Anthropic) at 65.5%.
- What's the best LMCA model you can run on a 24 GB GPU?
- Qwen3.8 27B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 15 GB), scoring 41.4% on LMCA.
- What's the best LMCA model you can run on a 12 GB GPU?
- Qwen3.5 9B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 5 GB), scoring 24.5% on LMCA.
- Can open models match proprietary models on LMCA?
- Not quite on LMCA: the strongest proprietary model (Claude Fable 5.1 (unspecified)) scores 65.5%, ahead of the best open model (GLM 5.3) at 55.5% — but you can run the open one yourself.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.