Reasoning

LMCA Leaderboard

LMCA (Language Model Conceptual Argumentation) scores models on expert-rated arguments across philosophy, decision theory, and AI risk — domains where there is no empirical feedback loop to check an answer against. It is part of the Conceptual Reasoning Index, built with Anthropic by researchers Emery Cooper and Caspar Oesterheld.

Source: epoch53 open models ranked+119 proprietaryData through Sep 2026

Open models ranked on LMCA

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 29GLM 5.3 · 753.3B
55.5%
2 / 38Kimi K3 · 2779.9B
52.7%
3 / 63DeepSeek V4.1 Flash · 763.2B
47.0%
4 / 67GLM 5.2 · 753.3B
45.8%
5 / 68DeepSeek V4 Pro 0813 · 1650.5B
45.5%
6 / 79DeepSeek V4 Flash 0731 · 304.2B
41.7%
7 / 81Qwen3.8 27B · 27.8B
41.4%
8 / 82DeepSeek V4 Pro · 1598.8B
41.2%
9 / 86Gemma 4 31B IT · 31.3B
39.3%
10 / 91Qwen3.5 397B A17B · 403.4B
37.9%
11 / 93Inkling · 952.4B
37.6%
12 / 96Kimi K2.6 · 1026.9B
37.3%
13 / 98NVIDIA Nemotron 3 Ultra 550B A55B BF16 · 560.5B
36.9%
14 / 101DeepSeek V4 Flash · 290.9B
35.9%
15 / 104Qwen3.6 27B · 27.8B
34.5%
16 / 106Qwen3.5 27B · 27.8B
34.0%
17 / 107MiniMax M3 · 427.0B
33.7%
18 / 109Qwen3.5 122B A10B · 125.1B
32.2%
19 / 112Qwen3.6 35B A3B · 36.0B
29.7%
20 / 113Gemma 4 26B A4B IT · 25.8B
29.7%
21 / 114MiMo V2.5 Pro · 1023.2B
29.5%
22 / 115Qwen3.5 35B A3B · 36.0B
29.5%
23 / 116Qwen3 235B A22B Thinking 2507 · 235.1B
29.3%
24 / 118DeepSeek V3.2 Exp · 685.4B
29.1%
25 / 120DeepSeek V3.2 · 685.4B
28.8%
26 / 121DeepSeek V3.1 Terminus · 684.5B
28.6%
27 / 127Qwen3 235B A22B · 235.1B
25.0%
28 / 128Qwen3.5 9B · 9.7B
24.5%
29 / 129DeepSeek V3.1 · 684.5B
24.3%
30 / 131Qwen3 235B A22B Instruct 2507 · 235.1B
23.5%
31 / 132Qwen3 30B A3B Instruct 2507 · 30.5B
22.4%
32 / 134GPT OSS 120B · 116.8B
22.1%
33 / 137Qwen3 30B A3B Thinking 2507 · 30.5B
19.5%
34 / 140Qwen3 14B · 14.8B
18.2%
35 / 142Llama 3.3 70B Instruct · 70.6B
17.5%
36 / 143Qwen3 32B · 32.8B
17.3%
37 / 146Mistral Large 3 675B Instruct 2512 · 675B
16.7%
38 / 148Llama 4 Maverick 17B 128E Instruct · 401.6B
15.9%
39 / 149Qwen3 30B A3B · 30.5B
15.8%
40 / 150DeepSeek v3 0324 · 684.5B
15.5%
41 / 152Llama 3.1 70B Instruct · 70.6B
14.8%
42 / 153GPT OSS 20B · 20.9B
14.5%
43 / 154Qwen2.5 72B Instruct · 72.7B
13.4%
44 / 155Gemma 3 27B IT · 27.4B
12.3%
45 / 156Llama 4 Scout 17B 16E Instruct · 108.6B
12.0%
46 / 158C4ai Command A 03 2025 · 111.1B
10.3%
47 / 161C4ai Command R 08 2024 · 32.3B
9.2%
48 / 163Qwen3 8B · 8.2B
8.8%
49 / 166Qwen2.5 7B Instruct · 7.6B
6.4%
50 / 169Llama 3.1 8B Instruct · 8.0B
5.4%
51 / 170C4ai Command R Plus 08 2024 · 103.8B
5.0%
52 / 171Gemma 3 12B IT · 12.2B
4.5%
53 / 172Gemma 3 4B IT · 4.3B
2.8%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

10B100B1Tmodel size (log scale) →55.5%2.8%Kimi K3 · 2.8T · 52.7%DeepSeek V4.1 Flash · 763B · 47.0%GLM 5.2 · 753B · 45.8%DeepSeek V4 Pro 0813 · 1.7T · 45.5%DeepSeek V4 Pro · 1.6T · 41.2%Gemma 4 31B IT · 31B · 39.3%Qwen3.5 397B A17B · 403B · 37.9%Inkling · 952B · 37.6%Kimi K2.6 · 1T · 37.3%NVIDIA Nemotron 3 Ultra 550B A55B BF16 · 561B · 36.9%DeepSeek V4 Flash · 291B · 35.9%Qwen3.6 27B · 28B · 34.5%Qwen3.5 27B · 28B · 34.0%MiniMax M3 · 427B · 33.7%Qwen3.5 122B A10B · 125B · 32.2%Qwen3.6 35B A3B · 36B · 29.7%MiMo V2.5 Pro · 1T · 29.5%Qwen3.5 35B A3B · 36B · 29.5%Qwen3 235B A22B Thinking 2507 · 235B · 29.3%DeepSeek V3.2 Exp · 685B · 29.1%DeepSeek V3.2 · 685B · 28.8%DeepSeek V3.1 Terminus · 685B · 28.6%Qwen3 235B A22B · 235B · 25.0%DeepSeek V3.1 · 685B · 24.3%Qwen3 235B A22B Instruct 2507 · 235B · 23.5%Qwen3 30B A3B Instruct 2507 · 31B · 22.4%GPT OSS 120B · 117B · 22.1%Qwen3 30B A3B Thinking 2507 · 31B · 19.5%Qwen3 14B · 15B · 18.2%Llama 3.3 70B Instruct · 71B · 17.5%Qwen3 32B · 33B · 17.3%Mistral Large 3 675B Instruct 2512 · 675B · 16.7%Llama 4 Maverick 17B 128E Instruct · 402B · 15.9%Qwen3 30B A3B · 31B · 15.8%DeepSeek v3 0324 · 685B · 15.5%Llama 3.1 70B Instruct · 71B · 14.8%GPT OSS 20B · 21B · 14.5%Qwen2.5 72B Instruct · 73B · 13.4%Gemma 3 27B IT · 27B · 12.3%Llama 4 Scout 17B 16E Instruct · 109B · 12.0%C4ai Command A 03 2025 · 111B · 10.3%C4ai Command R 08 2024 · 32B · 9.2%Llama 3.1 8B Instruct · 8B · 5.4%C4ai Command R Plus 08 2024 · 104B · 5.0%Gemma 3 12B IT · 12B · 4.5%Gemma 3 4B IT · 4B · 2.8%Gemma 3 4B ITQwen2.5 7B Instruct · 8B · 6.4%Qwen2.5 7B InstructQwen3 8B · 8B · 8.8%Qwen3 8BQwen3.5 9B · 10B · 24.5%Qwen3.5 9BGemma 4 26B A4B IT · 26B · 29.7%Gemma 4 26B A4B ITQwen3.8 27B · 28B · 41.4%Qwen3.8 27BDeepSeek V4 Flash 0731 · 304B · 41.7%DeepSeek V4 Flash 0731GLM 5.3 · 753B · 55.5%GLM 5.3
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Gemma 3 4B IT, 4B, score 2.8% — on the efficiency frontier (best score at its size or smaller).
  • Qwen2.5 7B Instruct, 8B, score 6.4% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3 8B, 8B, score 8.8% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.5 9B, 10B, score 24.5% — on the efficiency frontier (best score at its size or smaller).
  • Gemma 4 26B A4B IT, 26B, score 29.7% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.8 27B, 28B, score 41.4% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek V4 Flash 0731, 304B, score 41.7% — on the efficiency frontier (best score at its size or smaller).
  • GLM 5.3, 753B, score 55.5% — on the efficiency frontier (best score at its size or smaller).

LMCA: frequently asked questions

What is the best open LLM on LMCA?
GLM 5.3 is the top open model on LMCA, scoring 55.5%. Among all models tested — including proprietary ones — it ranks #29. The top model overall is Claude Fable 5.1 (unspecified) (Anthropic) at 65.5%.
What's the best LMCA model you can run on a 24 GB GPU?
Qwen3.8 27B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 15 GB), scoring 41.4% on LMCA.
What's the best LMCA model you can run on a 12 GB GPU?
Qwen3.5 9B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 5 GB), scoring 24.5% on LMCA.
Can open models match proprietary models on LMCA?
Not quite on LMCA: the strongest proprietary model (Claude Fable 5.1 (unspecified)) scores 65.5%, ahead of the best open model (GLM 5.3) at 55.5% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.