llmrun Composite Score · Reasoning

The best LLMs for reasoning, open and proprietary together, on one open-ended scale where 100 is the average model (reasoning). Switch to open models to see only what you can download and run.

40 benchmarks4 sourcesUpdated 3 Oct 2026Reference scale 2026-Q4How the Composite Score is built

Reasoning score vs VRAM needed

Proprietary models are not shown on this chart.

Models ranked by reasoning llmrun Composite Score
#Modelllmrun Composite ScoreCodingAgentsMathScienceReasoningBenchmarksVRAM
1GPT 6 Astra Max19 benchmarks15614415214216319–
2GPT 6.1 Sol Max13 benchmarks127–15314115713–
3Claude Opus 5.5 Max17 benchmarks147–15214415417–
4Claude Sonnet 5.5 Max13 benchmarks141–14614315113–
5GPT 5.6 Sol Promax6 benchmarks––––1486–
6GPT 5.6 Sol Max21 benchmarks14614114314214521–
7Claude Fable 5.1 Max15 benchmarks149–14914314315–
8GPT 6 Sol Max18 benchmarks149–14714114318–
9Claude Opus 5 Max22 benchmarks14614414313914322–
10GPT 5.4 Pro (Mar 05, 2026)8 benchmarks––1361411428–
11GPT 6 Astra6 benchmarks–164––1426–
12Claude Opus 5.54 benchmarks–––1451414–
13Claude Fable 5 Max20 benchmarks14713414813814120–
14Claude Fable 5.19 benchmarks147––1421419–
15Claude Opus 59 benchmarks147––1391419–
16Claude Fable 56 benchmarks––––1406–
17GPT 5.59 benchmarks141––1371399–
18GPT 5.6 Sol9 benchmarks142––1401399–
19GPT 5.6 Terra Max20 benchmarks13913814013913820–
20Gemini 3 Deep Think Preview3 benchmarks—––––1373–
21GPT 5.5 Pro7 benchmarks––142–1377–
22Gemini 3.7 Flash15 benchmarks118–13012913715–
23Gemini 3.1 Pro Preview8 benchmarks––––1358–
24GPT 6 Sol4 benchmarks–––1391344–
25GPT 5.4 (Mar 05, 2026)20 benchmarks13412913313413420–
26Claude Opus 4.8 Max19 benchmarks12712713612913319–
27Gemini 3.5 Flash4 benchmarks––––1324–
28GPT 5.6 Terra10 benchmarks135––13613210–
29Gemini 3.8 Flash6 benchmarks133––1251326–
30DeepSeek V4 Pro 0813open991 GB · 21 benchmarks125–12812613121991 GB
31Claude Opus 4.7 Max15 benchmarks131–13012013015–
32Gemini 3.6 Flash16 benchmarks113–12412513016–
33Kimi K3open1670 GB · 26 benchmarks139133135136129261670 GB
34Grok 4.79 benchmarks––1221301299–
35Grok 4.2011 benchmarks116122––12911–
36DeepSeek V4 Flash 0731open183 GB · 21 benchmarks121–12812412921183 GB
37Qwen3.8 Max10 benchmarks––132–12910–
38Muse Spark 1.27 benchmarks121––1311287–
39GPT 5.6 Luna Max20 benchmarks13413513513012820–
40GPT 5.2 (Dec 11, 2025)12 benchmarks––124–12812–
41Muse Spark 1.3 Max6 benchmarks––1321381286–
42Grok 4.516 benchmarks128–12512812716–
43Claude Sonnet 4.68 benchmarks–116––1278–
44GPT 6 Luna Max18 benchmarks134–13512812718–
45Claude Sonnet 4.6 Max11 benchmarks––12810812611–
46GLM 5.3open454 GB · 21 benchmarks13114012913112621454 GB
47Qwen3.7 Max17 benchmarks121–12612212517–
48GPT 5.6 Luna8 benchmarks–––1291258–
49Gemini 3 Flash Preview7 benchmarks––––1257–
50Claude Sonnet 5 Max13 benchmarks––13211412513–
51DeepSeek V4.1 Flashopen458 GB · 10 benchmarks116–13112412510458 GB
52Claude Opus 4.85 benchmarks–136––1245–
53Claude Opus 4.5 (Nov 01, 2025)9 benchmarks–125––1239–
54GPT 6 Luna5 benchmarks131––1271235–
55Claude Opus 4.6 Max12 benchmarks––12812412312–
56Claude Opus 4.77 benchmarks132147––1227–
57Inkling Smallopen163 GB · 12 benchmarks––11711812212163 GB
58Gemini 3 Pro Preview17 benchmarks12412712012612217–
59Inklingopen575 GB · 21 benchmarks104–11211612121575 GB
60Muse Spark 1.36 benchmarks––131–1216–
61Grok 4.311 benchmarks––11711812011–
62Qwen3.6 Max Preview8 benchmarks––––1208–
63GLM 5.2open454 GB · 24 benchmarks12013212612812024454 GB
64Kimi K2.6open618 GB · 18 benchmarks119–12212211918618 GB
65GPT 5 Pro (Oct 06, 2025)5 benchmarks––123–1195–
66Grok 4.1 Fast Reasoning6 benchmarks––––1196–
67Qwen3.8 27Bopen17.9 GB · 9 benchmarks112–118112119917.9 GB
68GPT 5.1 (Nov 13, 2025)14 benchmarks121–––11914–
69Kimi K2.7 Codeopen618 GB · 19 benchmarks11511912111811819618 GB
70Grok 4.20 0309 Reasoning8 benchmarks––118–1188–
71DeepSeek V4 Proopen960 GB · 22 benchmarks111–11812211722960 GB
72Grok 4 (Jul 09)12 benchmarks111103––11712–
73GPT 5 (Aug 07, 2025)20 benchmarks120–12211811720–
74GPT 5.4 Mini (Mar 17, 2026)12 benchmarks––12011811612–
75Claude Sonnet 4.5 (Sep 29, 2025)11 benchmarks–115–10511511–
76NVIDIA Nemotron 3 Ultra 550B A55B BF16open378 GB · 14 benchmarks103–11211211514378 GB
77Qwen3.7 Plus9 benchmarks103––1171159–
78GLM 5.1open454 GB · 13 benchmarks11412711711711513454 GB
79O3 (Apr 16, 2025)19 benchmarks114–11211011519–
80Claude Opus 4 (May 14, 2025)7 benchmarks–––961157–
81Claude Opus 4.1 (Aug 05, 2025)6 benchmarks––––1156–
82Qwen3.5 397B A17Bopen274 GB · 9 benchmarks––112–1159274 GB
83Qwen3.7 Flash5 benchmarks––109–1155–
84Kimi K2.5open618 GB · 20 benchmarks110111–11711420618 GB
85MiniMax M2.5open140 GB · 11 benchmarks109110––11411140 GB
86DeepSeek V3.2 Expopen413 GB · 10 benchmarks107––10111310413 GB
87O3 Pro (Jun 10, 2025)6 benchmarks117–––1136–
88Qwen3.5 122B A10Bopen88.1 GB · 6 benchmarks–––96112688.1 GB
89Gemini 3.5 Flash Lite11 benchmarks103–106–11211–
90DeepSeek V3.2open413 GB · 13 benchmarks95109––11213413 GB
91Qwen3.5 Flash10 benchmarks––106–11210–
92Gemini 2.5 Pro13 benchmarks–10210911211213–
93DeepSeek V4 Flashopen175 GB · 10 benchmarks107––11111210175 GB
94Qwen3.5 Plus11 benchmarks––––11211–
95GPT 5 Mini (Aug 07, 2025)15 benchmarks111–117–11115–
96MiMo V2.5 Proopen616 GB · 6 benchmarks116––1151116616 GB
97GPT 5.4 Nano (Mar 17, 2026)6 benchmarks–––1151116–
98O4 Mini (Apr 16, 2025)19 benchmarks110–11210811119–
99Qwen3.5 27Bopen17.9 GB · 6 benchmarks98–––111617.9 GB
100Claude 3.7 Sonnet (Feb 19, 2025)4 benchmarks––––1104–
101Grok 4 Fast6 benchmarks––––1106–
102Qwen3.6 Plus15 benchmarks108–11611411015–
103Gemma 4 31B ITopen21.7 GB · 11 benchmarks108––1051091121.7 GB
104GLM 5open454 GB · 15 benchmarks110118––10915454 GB
105Claude Haiku 4.5 (Oct 01, 2025)8 benchmarks––104–1098–
106DeepSeek V3.1 Terminusopen412 GB · 5 benchmarks108––1031085412 GB
107Tiny Recursion Model2 benchmarks—––––1072–
108MiniMax M3open258 GB · 19 benchmarks11311310611810719258 GB
109Gemini 3.1 Flash Lite6 benchmarks––109–1076–
110Qwen3.6 35B A3Bopen22.0 GB · 12 benchmarks105–1091101071222.0 GB
111Qwen3.6 27Bopen17.9 GB · 12 benchmarks102–1141111061217.9 GB
112Gemini 2.5 Pro Preview (Mar 25)6 benchmarks––––1066–
113GLM 5.3 Flashopen205 GB · 17 benchmarks11313512112410617205 GB
114Qwen Plus (Apr 28, 2025)2 benchmarks—––––1062–
115Qwen3 235B A22B Thinking 2507open143 GB · 13 benchmarks103––10810613143 GB
116Qwen3.5 35B A3Bopen22.0 GB · 8 benchmarks–––108106822.0 GB
117DeepSeek V3.1open412 GB · 7 benchmarks––––1057412 GB
118Qwen3 Max (Sep 23, 2025)13 benchmarks––105–10513–
119Gemini 2.5 Flash Preview (May 20)6 benchmarks––––1056–
120GLM 4.7open218 GB · 13 benchmarks10110611111110413218 GB

The range after each score is a 90% interval: given the boards a model has results on, its true score is very likely inside it. Short ranges mean many agreeing results; long ranges mean few or conflicting ones.

Why some models are missing. A model gets a score only with results on at least 4 benchmarks across at least 2 categories, one of them reasoning or coding. Skill scores need at least 2 benchmarks in that skill; otherwise they show as –.

VRAM is llmrun's estimate at Q4_K_M (or the smallest quantization we track) for the weights plus room for a 16K-token context. Longer contexts need more memory. Proprietary models can't run locally, so the VRAM filter hides them.

How the llmrun Composite Score is built · llmrun does not run these benchmarks; scores are aggregated from public sources.