Llama 3.1 Tulu 3 70B DPO — Benchmarks
Benchmark scores for Llama 3.1 Tulu 3 70B DPO aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.
No overall rank for Llama 3.1 Tulu 3 70B DPO: an overall score is only published when a model has been measured on enough of the benchmarks current models are still submitted to. Its individual scores below stand on their own — see the methodology for how the overall score is built.
Math
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| MATH Level 5 | 42.7% | #19 / 39 | #70 / 107 |
| AIME 2024/2025 | 4.4% | #67 / 82 | #242 / 270 |
Reasoning
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| GPQA Diamond | 46.3% | #60 / 94 | #231 / 289 |
Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.