Reasoning

LMCA Leaderboard

LMCA (Language Model Conceptual Argumentation) scores models on expert-rated arguments across philosophy, decision theory, and AI risk — domains where there is no empirical feedback loop to check an answer against. It is part of the Conceptual Reasoning Index, built with Anthropic by researchers Emery Cooper and Caspar Oesterheld.

Source: epoch53 open models ranked+119 proprietaryData through Sep 2026

All models ranked on LMCA

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1Claude Fable 5.1 (unspecified) · proprietary
65.5%
2Claude Fable 5.1 (medium) · proprietary
65.3%
3Claude Fable 5.1 (high) · proprietary
65.1%
4Claude Fable 5.1 (xhigh) · proprietary
65.1%
5Claude Opus 5 (xhigh) · proprietary
64.5%
6GPT 6 Astra (xhigh) · proprietary
64.4%
7GPT 6 Astra (unspecified) · proprietary
64.1%
8GPT 6 Astra (high) · proprietary
63.8%
9Claude Opus 5 (high) · proprietary
63.7%
10Claude Fable 5.1 (low) · proprietary
63.5%
11Claude Opus 5 Max · proprietary
63.3%
12Claude Opus 5 (medium) · proprietary
62.9%
13GPT 6 Astra (medium) · proprietary
62.9%
14Claude Opus 5 (low) · proprietary
62.6%
15GPT 6 Astra (low) · proprietary
62.2%
16Claude Fable 5 (high) · proprietary
61.1%
17Claude Fable 5 (xhigh) · proprietary
61.0%
18Claude Fable 5 Max · proprietary
60.3%
19Claude Fable 5 (medium) · proprietary
60.3%
20Claude Fable 5 (low) · proprietary
59.6%
21GPT 5.6 Sol Promax · proprietary
59.2%
22GPT 5.6 Sol (xhigh) · proprietary
58.5%
23GPT 5.6 Sol Max · proprietary
58.4%
24Claude Opus 4.8 Max · proprietary
57.5%
25GPT 5.6 Sol (high) · proprietary
56.9%
26GPT 5.6 Sol (medium) · proprietary
56.4%
27GPT 5.6 Sol (low) · proprietary
55.9%
28Claude Opus 4.6 Max · proprietary
55.8%
29GLM 5.3 · 753.3B
55.5%
30GPT 5.6 Terra Max · proprietary
55.0%
31GPT 5.5 (xhigh) · proprietary
54.3%
32GPT 5.5 Pro (xhigh) · proprietary
53.9%
33Muse Spark 1.3 (unspecified) · proprietary
53.9%
34Gemini 3.1 Pro Preview (high) · proprietary
53.8%
35Gemini 3.8 Flash (medium) · proprietary
52.9%
36Gemini 3.1 Pro Preview (medium) · proprietary
52.8%
37GPT 5.6 Sol None · proprietary
52.8%
38Kimi K3 · 2779.9B
52.7%
39GPT 5.6 Terra (xhigh) · proprietary
52.2%
40Claude Opus 4.7 Max · proprietary
52.2%
41GPT 5.4 (Mar 05, 2026, xhigh) · proprietary
52.0%
42Gemini 3.8 Flash (unspecified) · proprietary
51.5%
43GPT 5.6 Terra (high) · proprietary
51.4%
44Gemini 3.1 Pro Preview (low) · proprietary
51.2%
45Gemini 3.8 Flash (low) · proprietary
51.2%
46Gemini 3.7 Flash (high) · proprietary
50.4%
47Claude Sonnet 5 (high) · proprietary
50.0%
48Muse Spark 1.1 (high) · proprietary
49.9%
49Claude Sonnet 5 (medium) · proprietary
49.7%
50Claude Sonnet 5 (low) · proprietary
49.6%
51GPT 5.6 Terra (medium) · proprietary
49.6%
52Claude Sonnet 5 Max · proprietary
49.3%
53GPT 5.6 Terra (low) · proprietary
49.3%
54Gemini 3.7 Flash (medium) · proprietary
49.2%
55Claude Sonnet 5 (xhigh) · proprietary
48.8%
56GPT 5.6 Luna (xhigh) · proprietary
48.5%
57GPT 5.6 Luna Max · proprietary
48.5%
58Grok 4.6 (xhigh) · proprietary
48.5%
59Muse Spark 1.2 (xhigh) · proprietary
48.4%
60GPT 5.6 Luna (high) · proprietary
47.4%
61Gemini 3.5 Flash (high) · proprietary
47.1%
62Gemini 3.7 Flash (low) · proprietary
47.1%
63DeepSeek V4.1 Flash · 763.2B
47.0%
64Claude Sonnet 4.6 Max · proprietary
46.5%
65GPT 5.6 Terra None · proprietary
46.5%
66Qwen3.8 Max (xhigh) · proprietary
46.2%
67GLM 5.2 · 753.3B
45.8%
68DeepSeek V4 Pro 0813 · 1650.5B
45.5%
69Grok 4.5 (high) · proprietary
45.2%
70Gemini 3.6 Flash (high) · proprietary
44.9%
71Claude Opus 4.5 (Nov 01, 2025, high) · proprietary
44.5%
72GPT 5.6 Luna (medium) · proprietary
44.3%
73Qwen3.7 Max · proprietary
44.0%
74GPT 5.2 (Dec 11, 2025, xhigh) · proprietary
43.9%
75GPT 5.1 (Nov 13, 2025, high) · proprietary
43.9%
76GPT 5.6 Luna (low) · proprietary
43.8%
77Gemini 3 Flash Preview (high) · proprietary
43.1%
78Qwen3.6 Max Preview · proprietary
42.5%
79DeepSeek V4 Flash 0731 · 304.2B
41.7%
80GPT 5.6 Luna None · proprietary
41.4%
81Qwen3.8 27B · 27.8B
41.4%
82DeepSeek V4 Pro · 1598.8B
41.2%
83GPT 5.4 Mini (Mar 17, 2026, xhigh) · proprietary
40.8%
84GPT 5 (Aug 07, 2025, high) · proprietary
40.0%
85O3 (Apr 16, 2025, high) · proprietary
39.7%
86Gemma 4 31B IT · 31.3B
39.3%
87Claude Sonnet 4.5 (Sep 29, 2025, unspecified) · proprietary
38.8%
88Grok 4.20 · proprietary
38.7%
89O3 Pro (Jun 10, 2025, high) · proprietary
38.5%
90Grok 4.3 (high) · proprietary
38.3%
91Qwen3.5 397B A17B · 403.4B
37.9%
92Gemini 3.5 Flash Lite (high) · proprietary
37.6%
93Inkling · 952.4B
37.6%
94Qwen3.7 Plus · proprietary
37.6%
95Claude Opus 4 (May 14, 2025, unspecified) · proprietary
37.4%
96Kimi K2.6 · 1026.9B
37.3%
97Claude Opus 4.1 (Aug 05, 2025, unspecified) · proprietary
37.1%
98NVIDIA Nemotron 3 Ultra 550B A55B BF16 · 560.5B
36.9%
99GPT 5.4 Nano (Mar 17, 2026, xhigh) · proprietary
36.9%
100Qwen3.5 Plus · proprietary
36.4%
101DeepSeek V4 Flash · 290.9B
35.9%
102Gemini 3.1 Flash Lite (high) · proprietary
35.0%
103Gemini 2.5 Pro · proprietary
34.8%
104Qwen3.6 27B · 27.8B
34.5%
105GPT 5 Mini (Aug 07, 2025, high) · proprietary
34.2%
106Qwen3.5 27B · 27.8B
34.0%
107MiniMax M3 · 427.0B
33.7%
108Qwen3.6 Plus · proprietary
33.1%
109Qwen3.5 122B A10B · 125.1B
32.2%
110Qwen3.6 Flash · proprietary
31.0%
111Claude Haiku 4.5 (Oct 01, 2025, unspecified) · proprietary
30.9%
112Qwen3.6 35B A3B · 36.0B
29.7%
113Gemma 4 26B A4B IT · 25.8B
29.7%
114MiMo V2.5 Pro · 1023.2B
29.5%
115Qwen3.5 35B A3B · 36.0B
29.5%
116Qwen3 235B A22B Thinking 2507 · 235.1B
29.3%
117Qwen3.5 Flash · proprietary
29.1%
118DeepSeek V3.2 Exp · 685.4B
29.1%
119Claude Sonnet 4 (May 14, 2025, unspecified) · proprietary
29.0%
120DeepSeek V3.2 · 685.4B
28.8%
121DeepSeek V3.1 Terminus · 684.5B
28.6%
122Qwen3 Max (Sep 23, 2025) · proprietary
28.3%
123Gemini 2.5 Flash · proprietary
27.5%
124O4 Mini (Apr 16, 2025, high) · proprietary
26.5%
125Mistral Medium 2604 · proprietary
26.1%
126GPT 4.1 (Apr 14, 2025) · proprietary
25.6%
127Qwen3 235B A22B · 235.1B
25.0%
128Qwen3.5 9B · 9.7B
24.5%
129DeepSeek V3.1 · 684.5B
24.3%
130Qwen Plus (Apr 28, 2025) · proprietary
24.0%
131Qwen3 235B A22B Instruct 2507 · 235.1B
23.5%
132Qwen3 30B A3B Instruct 2507 · 30.5B
22.4%
133O1 (Dec 17, 2024, high) · proprietary
22.3%
134GPT OSS 120B · 116.8B
22.1%
135GPT 4.1 Mini (Apr 14, 2025) · proprietary
21.1%
136Mistral Small 2603 · proprietary
20.6%
137Qwen3 30B A3B Thinking 2507 · 30.5B
19.5%
138Mistral Medium 2508 · proprietary
19.1%
139O3 Mini (Jan 31, 2025, high) · proprietary
19.0%
140Qwen3 14B · 14.8B
18.2%
141Gemini 2.5 Flash Lite Preview Thinking (Jun 17) · proprietary
18.1%
142Llama 3.3 70B Instruct · 70.6B
17.5%
143Qwen3 32B · 32.8B
17.3%
144GPT 4 (Jun 13) · proprietary
17.1%
145Claude 3 Opus (Feb 29, 2024) · proprietary
17.0%
146Mistral Large 3 675B Instruct 2512 · 675B
16.7%
147GPT 4o (May 13, 2024) · proprietary
16.6%
148Llama 4 Maverick 17B 128E Instruct · 401.6B
15.9%
149Qwen3 30B A3B · 30.5B
15.8%
150DeepSeek v3 0324 · 684.5B
15.5%
151DeepSeek Chat · proprietary
15.2%
152Llama 3.1 70B Instruct · 70.6B
14.8%
153GPT OSS 20B · 20.9B
14.5%
154Qwen2.5 72B Instruct · 72.7B
13.4%
155Gemma 3 27B IT · 27.4B
12.3%
156Llama 4 Scout 17B 16E Instruct · 108.6B
12.0%
157GPT 4o Mini (Jul 18, 2024) · proprietary
10.4%
158C4ai Command A 03 2025 · 111.1B
10.3%
159GPT 4 Turbo · proprietary
9.8%
160GPT 3.5 Turbo (Jan 25) · proprietary
9.7%
161C4ai Command R 08 2024 · 32.3B
9.2%
162Claude 3 Haiku (Mar 07, 2024) · proprietary
8.8%
163Qwen3 8B · 8.2B
8.8%
164GPT 5 Nano (Aug 07, 2025, high) · proprietary
7.9%
165Gemma 2 27B · proprietary
7.1%
166Qwen2.5 7B Instruct · 7.6B
6.4%
167Ministral 3B 2410 · proprietary
5.5%
168GPT 4.1 Nano (Apr 14, 2025) · proprietary
5.5%
169Llama 3.1 8B Instruct · 8.0B
5.4%
170C4ai Command R Plus 08 2024 · 103.8B
5.0%
171Gemma 3 12B IT · 12.2B
4.5%
172Gemma 3 4B IT · 4.3B
2.8%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

10B100B1Tmodel size (log scale) →55.5%2.8%Kimi K3 · 2.8T · 52.7%DeepSeek V4.1 Flash · 763B · 47.0%GLM 5.2 · 753B · 45.8%DeepSeek V4 Pro 0813 · 1.7T · 45.5%DeepSeek V4 Pro · 1.6T · 41.2%Gemma 4 31B IT · 31B · 39.3%Qwen3.5 397B A17B · 403B · 37.9%Inkling · 952B · 37.6%Kimi K2.6 · 1T · 37.3%NVIDIA Nemotron 3 Ultra 550B A55B BF16 · 561B · 36.9%DeepSeek V4 Flash · 291B · 35.9%Qwen3.6 27B · 28B · 34.5%Qwen3.5 27B · 28B · 34.0%MiniMax M3 · 427B · 33.7%Qwen3.5 122B A10B · 125B · 32.2%Qwen3.6 35B A3B · 36B · 29.7%MiMo V2.5 Pro · 1T · 29.5%Qwen3.5 35B A3B · 36B · 29.5%Qwen3 235B A22B Thinking 2507 · 235B · 29.3%DeepSeek V3.2 Exp · 685B · 29.1%DeepSeek V3.2 · 685B · 28.8%DeepSeek V3.1 Terminus · 685B · 28.6%Qwen3 235B A22B · 235B · 25.0%DeepSeek V3.1 · 685B · 24.3%Qwen3 235B A22B Instruct 2507 · 235B · 23.5%Qwen3 30B A3B Instruct 2507 · 31B · 22.4%GPT OSS 120B · 117B · 22.1%Qwen3 30B A3B Thinking 2507 · 31B · 19.5%Qwen3 14B · 15B · 18.2%Llama 3.3 70B Instruct · 71B · 17.5%Qwen3 32B · 33B · 17.3%Mistral Large 3 675B Instruct 2512 · 675B · 16.7%Llama 4 Maverick 17B 128E Instruct · 402B · 15.9%Qwen3 30B A3B · 31B · 15.8%DeepSeek v3 0324 · 685B · 15.5%Llama 3.1 70B Instruct · 71B · 14.8%GPT OSS 20B · 21B · 14.5%Qwen2.5 72B Instruct · 73B · 13.4%Gemma 3 27B IT · 27B · 12.3%Llama 4 Scout 17B 16E Instruct · 109B · 12.0%C4ai Command A 03 2025 · 111B · 10.3%C4ai Command R 08 2024 · 32B · 9.2%Llama 3.1 8B Instruct · 8B · 5.4%C4ai Command R Plus 08 2024 · 104B · 5.0%Gemma 3 12B IT · 12B · 4.5%Gemma 3 4B IT · 4B · 2.8%Gemma 3 4B ITQwen2.5 7B Instruct · 8B · 6.4%Qwen2.5 7B InstructQwen3 8B · 8B · 8.8%Qwen3 8BQwen3.5 9B · 10B · 24.5%Qwen3.5 9BGemma 4 26B A4B IT · 26B · 29.7%Gemma 4 26B A4B ITQwen3.8 27B · 28B · 41.4%Qwen3.8 27BDeepSeek V4 Flash 0731 · 304B · 41.7%DeepSeek V4 Flash 0731GLM 5.3 · 753B · 55.5%GLM 5.3
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Gemma 3 4B IT, 4B, score 2.8% — on the efficiency frontier (best score at its size or smaller).
  • Qwen2.5 7B Instruct, 8B, score 6.4% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3 8B, 8B, score 8.8% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.5 9B, 10B, score 24.5% — on the efficiency frontier (best score at its size or smaller).
  • Gemma 4 26B A4B IT, 26B, score 29.7% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.8 27B, 28B, score 41.4% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek V4 Flash 0731, 304B, score 41.7% — on the efficiency frontier (best score at its size or smaller).
  • GLM 5.3, 753B, score 55.5% — on the efficiency frontier (best score at its size or smaller).

LMCA: frequently asked questions

What is the best open LLM on LMCA?
GLM 5.3 is the top open model on LMCA, scoring 55.5%. Among all models tested — including proprietary ones — it ranks #29. The top model overall is Claude Fable 5.1 (unspecified) (Anthropic) at 65.5%.
What's the best LMCA model you can run on a 24 GB GPU?
Qwen3.8 27B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 15 GB), scoring 41.4% on LMCA.
What's the best LMCA model you can run on a 12 GB GPU?
Qwen3.5 9B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 5 GB), scoring 24.5% on LMCA.
Can open models match proprietary models on LMCA?
Not quite on LMCA: the strongest proprietary model (Claude Fable 5.1 (unspecified)) scores 65.5%, ahead of the best open model (GLM 5.3) at 55.5% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.