Phi 4 Reasoning — Benchmarks

Benchmark scores for Phi 4 Reasoning aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.

No overall rank for Phi 4 Reasoning: an overall score is only published when a model has been measured on enough of the benchmarks current models are still submitted to. Its individual scores below stand on their own — see the methodology for how the overall score is built.

Knowledge

BenchmarkScoreOpen rankAll models
MMLU-Pro74.3%#32 / 144#88 / 257

Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.