Phi 3 Medium 128k Instruct — Benchmarks

Benchmark scores for Phi 3 Medium 128k Instruct aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.

No overall rank for Phi 3 Medium 128k Instruct: an overall score is only published when a model has been measured on enough of the benchmarks current models are still submitted to. Its individual scores below stand on their own — see the methodology for how the overall score is built.

Knowledge

BenchmarkScoreOpen rankAll models
MMLU-Pro51.9%#73 / 144#164 / 257
ARC Challenge91.6%#5 / 56#5 / 76
OpenBookQA87.4%#3 / 26#3 / 41
TriviaQA73.9%#9 / 19#24 / 39
MMLU78.0%#16 / 85#34 / 136
HellaSwag82.4%#12 / 47#21 / 75

Math

BenchmarkScoreOpen rankAll models
MATH Level 517.6%#32 / 39#91 / 107

Reasoning

BenchmarkScoreOpen rankAll models
GPQA Diamond27.6%#83 / 94#272 / 289
WinoGrande81.5%#8 / 51#16 / 79
BIG-Bench Hard81.4%#3 / 38#5 / 50

Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.