Best LLMs for math: open and proprietary

Open and proprietary models together on one 0–100 scale (mathematics). Switch to local models to see only what you can download and run.

40 benchmarks4 sourcesUpdated 3 Oct 2026Reference scale 2026-Q4How the score is built

Models ranked by math llmrun Score
#Modelllmrun ScoreCodingAgentsMathScienceReasoningBenchmarksVRAM
121Gemini 1.5 Pro 0015 benchmarks––19–205–
122Llama 3.2 90B Vision Instructopen58.5 GB · 4 benchmarks––18––458.5 GB
123Claude 3 Opus (Feb 29, 2024)10 benchmarks––18–2210–
124Llama 3.1 70B Instructopen46.6 GB · 8 benchmarks––18–22846.6 GB
125Gemma 2 27B ITopen18.0 GB · 4 benchmarks––16––418.0 GB
126Gemini 1.5 Flash 0015 benchmarks––15–165–
127Mistral Large 24024 benchmarks––15––4–
128Meta Llama 3 70B Instructopen46.6 GB · 5 benchmarks––14––546.6 GB
129GPT 4 (Jun 13)9 benchmarks––14–239–
130Hermes 2 Theta Llama 3 70Bopen43.3 GB · 3 benchmarks—––14––343.3 GB
131Llama 3.1 8B Instructopen5.3 GB · 11 benchmarks14–14716115.3 GB
132Gemma 2 9B ITopen6.1 GB · 4 benchmarks––14––46.1 GB
133Claude 3 Sonnet (Feb 29, 2024)5 benchmarks––13––5–
134Claude 3 Haiku (Mar 07, 2024)7 benchmarks––12–157–
135Claude 2.04 benchmarks––12––4–
136GPT 3.5 Turbo (Jan 25)10 benchmarks––12–1210–
137Gemini 1.0 Pro 0014 benchmarks––11––4–
138Deepseek Llm 67B Chatopen41.3 GB · 4 benchmarks––10––441.3 GB
139Meta Llama 3 8B Instructopen5.3 GB · 7 benchmarks––10–1475.3 GB
140Mistral 7B Instruct v0.3open4.9 GB · 4 benchmarks––10––44.9 GB
141Llama 2 70B Chat HFopen45.5 GB · 5 benchmarks––9–10545.5 GB

The range after each score is a 90% interval: given the boards a model has results on, its true score is very likely inside it. Short ranges mean many agreeing results; long ranges mean few or conflicting ones.

Why some models are missing. A model gets a score only with results on at least 4 benchmarks across at least 2 categories, one of them reasoning or coding. Skill scores need at least 2 benchmarks in that skill; otherwise they show as –.

VRAM is llmrun's estimate at Q4_K_M (or the smallest quantization we track) for the weights plus a working context. Proprietary models can't run locally, so the VRAM filter hides them.

How the llmrun Score is built · llmrun does not run these benchmarks; scores are aggregated from public sources.