Llama 4 Maverick 17B 128E Instruct — Benchmarks

Benchmark scores for Llama 4 Maverick 17B 128E Instruct aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.

Overall rank: #59 of 78 open modelscomposite 33.3/100 across 14 benchmarks in 4 categories · methodology

Coding

BenchmarkScoreOpen rankAll models
SciCode33.1%#46 / 62#141 / 161
Aider Polyglot15.6%#15 / 18#62 / 69
WeirdML24.5%#39 / 46#143 / 161

Instruction Following

BenchmarkScoreOpen rankAll models
LMArena Text1326.9#85 / 182#226 / 391

Knowledge

BenchmarkScoreOpen rankAll models
Humanity's Last Exam5.7%#4 / 4#41 / 49
MMLU-Pro80.5%#24 / 144#62 / 257

Math

BenchmarkScoreOpen rankAll models
MATH Level 573.0%#7 / 39#42 / 107
AIME 2024/202520.6%#52 / 82#212 / 270

Reasoning

BenchmarkScoreOpen rankAll models
DTBench61.9%#40 / 67#158 / 210
LMCA15.9%#38 / 53#148 / 172
GPQA Diamond67.0%#34 / 94#172 / 289
SimpleBench27.7%#17 / 25#82 / 101
CritPt0.0%#51 / 62#155 / 172
ARC-AGI4.4%#15 / 16#195 / 200
ARC-AGI-20.0%#14 / 16#187 / 190
Fiction.LiveBench46.2%#18 / 24#48 / 58

Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.