Phi 4 — Benchmarks

Benchmark scores for Phi 4 aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.

Overall rank: #52 of 123 open modelscomposite 51.4/100 across 5 benchmarks in 3 categories · methodology

Knowledge

BenchmarkScoreOpen rankAll models
MMLU-Pro70.4%#35 / 119#99 / 259
MMLU84.8%#5 / 77#11 / 136

Math

BenchmarkScoreOpen rankAll models
AIME 2024/202513.8%#51 / 74#216 / 270
MATH Level 564.9%#9 / 32#49 / 108

Reasoning

BenchmarkScoreOpen rankAll models
GPQA Diamond56.1%#42 / 83#200 / 291

Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.