Llama 3.1 Tulu 3 70B DPO — Benchmarks

Benchmark scores for Llama 3.1 Tulu 3 70B DPO aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.

No overall rank for Llama 3.1 Tulu 3 70B DPO: an overall score is only published when a model has been measured on enough of the benchmarks current models are still submitted to. Its individual scores below stand on their own — see the methodology for how the overall score is built.

Math

BenchmarkScoreOpen rankAll models
MATH Level 542.7%#19 / 39#70 / 107
AIME 2024/20254.4%#67 / 82#242 / 270

Reasoning

BenchmarkScoreOpen rankAll models
GPQA Diamond46.3%#60 / 94#231 / 289

Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.