Phi 3 Medium 128k Instruct — Benchmarks
Benchmark scores for Phi 3 Medium 128k Instruct aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.
No overall rank for Phi 3 Medium 128k Instruct: an overall score is only published when a model has been measured on enough of the benchmarks current models are still submitted to. Its individual scores below stand on their own — see the methodology for how the overall score is built.
Knowledge
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| MMLU-Pro | 51.9% | #73 / 144 | #164 / 257 |
| ARC Challenge | 91.6% | #5 / 56 | #5 / 76 |
| OpenBookQA | 87.4% | #3 / 26 | #3 / 41 |
| TriviaQA | 73.9% | #9 / 19 | #24 / 39 |
| MMLU | 78.0% | #16 / 85 | #34 / 136 |
| HellaSwag | 82.4% | #12 / 47 | #21 / 75 |
Math
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| MATH Level 5 | 17.6% | #32 / 39 | #91 / 107 |
Reasoning
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| GPQA Diamond | 27.6% | #83 / 94 | #272 / 289 |
| WinoGrande | 81.5% | #8 / 51 | #16 / 79 |
| BIG-Bench Hard | 81.4% | #3 / 38 | #5 / 50 |
Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.