llmrun Composite Score · Reasoning

The best LLMs for reasoning, open and proprietary together, on one open-ended scale where 100 is the average model (reasoning). Switch to open models to see only what you can download and run.

40 benchmarks4 sourcesUpdated 3 Oct 2026Reference scale 2026-Q4How the Composite Score is built

Reasoning score vs VRAM needed

Proprietary models are not shown on this chart.

Models ranked by reasoning llmrun Composite Score
#Modelllmrun Composite ScoreCodingAgentsMathScienceReasoningBenchmarksVRAM
121Qwen3.6 Flash10 benchmarks––108–10410–
122O1 (Dec 17, 2024)8 benchmarks––104–1038–
123Gemini 2.5 Flash9 benchmarks–94––1039–
124Codex Mini (May 16, 2025)2 benchmarks—––––1022–
125Mistral Medium 26049 benchmarks102––1011029–
126Gemma 4 26B A4B ITopen16.5 GB · 9 benchmarks104––102101916.5 GB
127GPT 5.4 2026 03.056 benchmarks118–––1016–
128DeepSeek R1 0528open412 GB · 11 benchmarks107–104–10011412 GB
129O3 Mini (Jan 31, 2025)18 benchmarks103–1061059918–
130Qwen3 235B A22Bopen143 GB · 10 benchmarks101–––9910143 GB
131O1 Preview (Sep 12, 2024)6 benchmarks––91–996–
132Qwen3 235B A22B Instruct 2507open143 GB · 8 benchmarks102–––988143 GB
133Qwen3.5 9Bopen6.6 GB · 9 benchmarks–––1049796.6 GB
134Grok 3 Mini2 benchmarks—––––972–
135GPT OSS 120Bopen70.8 GB · 20 benchmarks9882–103972070.8 GB
136Claude 3.5 Sonnet (Oct 22, 2024)11 benchmarks9910081889611–
137DeepSeek R1open412 GB · 14 benchmarks100–99999614412 GB
138GPT 5.4 Mini 2026 03.176 benchmarks––89–966–
139Mistral Small 4 119B 2603open81.6 GB · 3 benchmarks—––––96381.6 GB
140GPT 4.5 Preview (Feb 27, 2025)11 benchmarks96–92989511–
141GPT 5.2 2025 12.116 benchmarks––––956–
142GPT 5.4 Nano 2026 03.176 benchmarks––95–946–
143Qwen3 30B A3B Thinking 2507open19.4 GB · 8 benchmarks–––9994819.4 GB
144GPT 5.1 2025 11.137 benchmarks–––96937–
145Qwen3 30B A3B Instruct 2507open19.4 GB · 5 benchmarks––––92519.4 GB
146Mistral Medium 25086 benchmarks92––93926–
147GPT 4.1 (Apr 14, 2025)17 benchmarks99–93969217–
148Grok 34 benchmarks––––924–
149Claude 3.5 Sonnet (Jun 20, 2024)7 benchmarks––79–927–
150GPT 4.1 Mini (Apr 14, 2025)16 benchmarks92–95979116–
151Qwen3 32Bopen22.6 GB · 9 benchmarks93––9691922.6 GB
152GPT OSS 20Bopen13.1 GB · 11 benchmarks101––92911113.1 GB
153GPT 5 Nano (Aug 07, 2025)15 benchmarks102–108–9015–
154Kimi K2 Instructopen617 GB · 10 benchmarks102102––9010617 GB
155Magistral Medium 25062 benchmarks—––––892–
156Mistral Large 3 675B Instruct 2512open454 GB · 5 benchmarks93––96895454 GB
157Magistral Small 2506open17.8 GB · 4 benchmarks––––88417.8 GB
158GPT 4o (May 13, 2024)7 benchmarks––79–887–
159Qwen3 14Bopen11.9 GB · 8 benchmarks–––9488811.9 GB
160DeepSeek v3 0324open412 GB · 13 benchmarks99–91978813412 GB
161GPT 4o (Nov 20, 2024)10 benchmarks82–78828810–
162Gemini 2.0 Flash 0017 benchmarks––91–877–
163Grok 2 (Dec 12)6 benchmarks––83–876–
164Gemini 2.5 Flash Lite Preview Thinking (Jun 17)5 benchmarks95–––875–
165GPT 4 (Jun 13)9 benchmarks––68–879–
166GPT 4 Turbo (Apr 09, 2024)8 benchmarks––77–878–
167DeepSeek Chat6 benchmarks––––866–
168Qwen2.5 72B Instructopen49.3 GB · 9 benchmarks––82–86949.3 GB
169Llama 4 Maverick 17B 128E Instructopen273 GB · 17 benchmarks81–87968517273 GB
170Llama 3.1 405B Instructopen276 GB · 8 benchmarks––79–858276 GB
171Magistral Small 2509open18.1 GB · 6 benchmarks–––8385618.1 GB
172O1 Mini (Sep 12, 2024)6 benchmarks––95–856–
173Claude 3 Opus (Feb 29, 2024)10 benchmarks––74–8410–
174Mistral Large Instruct 2411open79.8 GB · 6 benchmarks––79–84679.8 GB
175GPT 4 Turbo2 benchmarks—––––842–
176C4ai Command A 03 2025open78.7 GB · 3 benchmarks—––––83378.7 GB
177Qwen3 30B A3Bopen19.4 GB · 7 benchmarks––––83719.4 GB
178Llama 3.1 70B Instructopen51.3 GB · 8 benchmarks––74–83851.3 GB
179Mistral Large Instruct 2407open86.4 GB · 7 benchmarks––77–83786.4 GB
180Mistral Small 3.2 24B Instruct 2506open18.1 GB · 6 benchmarks–––8283618.1 GB
181Gemini 1.5 Pro 0029 benchmarks––8789839–
182Llama 3.3 70B Instructopen51.3 GB · 13 benchmarks75–7681821351.3 GB
183Mistral Small 3.1 24B Instruct 2503open18.1 GB · 7 benchmarks––778182718.1 GB
184DeepSeek v3open412 GB · 10 benchmarks97–84898110412 GB
185Qwen3 8Bopen7.6 GB · 8 benchmarks–––888187.6 GB
186Gemini 1.5 Pro 0015 benchmarks––76–805–
187Llama 4 Scout 17B 16E Instructopen77.0 GB · 12 benchmarks––8283801277.0 GB
188GPT 4o (Aug 06, 2024)8 benchmarks––79–788–
189Mixtral 8x22B Instruct v0.1open88.4 GB · 6 benchmarks––––76688.4 GB
190C4ai Command R Plus 08 2024open73.8 GB · 3 benchmarks—––––74373.8 GB
191Gemma 3 27B ITopen21.7 GB · 11 benchmarks69–8780741121.7 GB
192GPT 4o Mini (Jul 18, 2024)14 benchmarks68–79–7214–
193GPT 4.1 Nano (Apr 14, 2025)13 benchmarks75–88827213–
194Ministral 3B 24104 benchmarks––––714–
195Llama 3.1 8B Instructopen7.9 GB · 11 benchmarks56–685569117.9 GB
196Mixtral 8x7B Instruct v0.1open30.5 GB · 5 benchmarks––––69530.5 GB
197Gemma 3 4B ITopen5.0 GB · 6 benchmarks––––6865.0 GB
198Gemini 1.5 Flash 0015 benchmarks––69–685–
199Claude 2.15 benchmarks––––685–
200Gemma 3 12B ITopen10.9 GB · 8 benchmarks–––7266810.9 GB
201Gemma 2 27B2 benchmarks—––––662–
202Qwen2.5 7B Instructopen5.8 GB · 6 benchmarks––––6565.8 GB
203Claude 3 Haiku (Mar 07, 2024)7 benchmarks––64–657–
204C4ai Command R 08 2024open25.1 GB · 2 benchmarks—––––64225.1 GB
205Meta Llama 3 8B Instructopen7.9 GB · 7 benchmarks––59–6477.9 GB
206Llama 2 13B Chat HFopen11.5 GB · 2 benchmarks—––––57211.5 GB
207GPT 3.5 Turbo (Jan 25)10 benchmarks––62–5710–
208Llama 2 70B Chat HFopen50.2 GB · 5 benchmarks––56–45550.2 GB

The range after each score is a 90% interval: given the boards a model has results on, its true score is very likely inside it. Short ranges mean many agreeing results; long ranges mean few or conflicting ones.

Why some models are missing. A model gets a score only with results on at least 4 benchmarks across at least 2 categories, one of them reasoning or coding. Skill scores need at least 2 benchmarks in that skill; otherwise they show as –.

VRAM is llmrun's estimate at Q4_K_M (or the smallest quantization we track) for the weights plus room for a 16K-token context. Longer contexts need more memory. Proprietary models can't run locally, so the VRAM filter hides them.

How the llmrun Composite Score is built · llmrun does not run these benchmarks; scores are aggregated from public sources.