Best LLMs for coding: open and proprietary

Open and proprietary models together on one 0–100 scale (coding). Switch to local models to see only what you can download and run.

40 benchmarks4 sourcesUpdated 3 Oct 2026Reference scale 2026-Q4How the score is built

Models ranked by coding llmrun Score
#Modelllmrun ScoreCodingAgentsMathScienceReasoningBenchmarksVRAM
121Gemma 3 27B ITopen18.1 GB · 11 benchmarks18–2916171118.1 GB
122GPT 4o Mini (Jul 18, 2024)14 benchmarks18–22–1714–
123Llama 3.1 8B Instructopen5.3 GB · 11 benchmarks14–14716115.3 GB

The range after each score is a 90% interval: given the boards a model has results on, its true score is very likely inside it. Short ranges mean many agreeing results; long ranges mean few or conflicting ones.

Why some models are missing. A model gets a score only with results on at least 4 benchmarks across at least 2 categories, one of them reasoning or coding. Skill scores need at least 2 benchmarks in that skill; otherwise they show as –.

VRAM is llmrun's estimate at Q4_K_M (or the smallest quantization we track) for the weights plus a working context. Proprietary models can't run locally, so the VRAM filter hides them.

How the llmrun Score is built · llmrun does not run these benchmarks; scores are aggregated from public sources.