Reasoning

BALROG Leaderboard

BALROG (Benchmarking Agentic LLM and VLM Reasoning On Games) tests long-horizon agentic reasoning across six game environments — BabyAI, Crafter, TextWorld, Baba Is AI, MiniHack and NetHack — scoring the average progress a model makes toward completing each game. It was built by UCL's DARK lab and collaborators.

Source: epoch14 open models ranked+26 proprietaryData through Sep 2026

All models ranked on BALROG

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1GPT 6 Astra Max · proprietary
68.3%
2Claude Opus 5 Max · proprietary
63.4%
3GPT 5.6 Sol Max · proprietary
60.0%
4Gemini 3 Pro Preview · proprietary
58.1%
5Gemini 3.1 Pro Preview · proprietary
57.0%
6GPT 5.6 Terra Max · proprietary
53.2%
7Gemini 3 Flash Preview · proprietary
48.1%
8GPT 5.6 Luna Max · proprietary
45.6%
9Grok 4 (Jul 09) · proprietary
43.6%
10Claude Opus 4.5 (Nov 01, 2025) · proprietary
43.5%
11Gemini 2.5 Pro Exp (Mar 25) · proprietary
43.3%
12Claude Opus 4.5 (Nov 01, 2025, 64K) · proprietary
43.0%
13Claude Opus 4.5 (Nov 01, 2025, unspecified) · proprietary
43.0%
14DeepSeek R1 · 684.5B
34.9%
15Gemini 2.5 Flash · proprietary
33.5%
16GPT 5 (Aug 07, 2025, minimal) · proprietary
32.8%
17Claude 3.5 Sonnet (Oct 22, 2024) · proprietary
32.6%
18GPT 4o (May 13, 2024) · proprietary
32.3%
19Claude Haiku 4.5 (Oct 01, 2025, 1K) · proprietary
31.2%
20Claude Haiku 4.5 (Oct 01, 2025) · proprietary
31.2%
21Grok 3 Beta · proprietary
29.5%
22Reka Flash 3 · proprietary
29.2%
23Llama 3.1 70B Instruct · 70.6B
27.9%
24Llama 3.2 90B Vision Instruct · 88.6B
27.3%
25Llama 3.3 70B Instruct · 70.6B
23.0%
26Gemini 1.5 Pro 002 · proprietary
21.0%
27DeepSeek R1 Distill Qwen 32B · 32.8B
19.5%
28Claude 3.5 Haiku (Oct 22, 2024) · proprietary
19.3%
29Mistral Nemo Instruct 2407 · 12.2B
17.6%
30GPT 4o Mini (Jul 18, 2024) · proprietary
17.4%
31Llama 3.2 11B Vision Instruct · 10.7B
16.8%
32Qwen2.5 72B Instruct · 72.7B
16.2%
33Llama 3.1 8B Instruct · 8.0B
15.1%
34Gemini 1.5 Flash 002 · proprietary
14.6%
35Qwen2 VL 72B Instruct · proprietary
12.8%
36Phi 4 · 14.7B
11.6%
37Llama 3.2 3B Instruct · 3.2B
10.1%
38Qwen2.5 7B Instruct · 7.6B
7.8%
39Llama 3.2 1B Instruct · 1.2B
6.6%
40Qwen2 VL 7B Instruct · 8.3B
3.7%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

10B100Bmodel size (log scale) →34.9%3.7%Llama 3.2 90B Vision Instruct · 89B · 27.3%Llama 3.3 70B Instruct · 71B · 23.0%Qwen2.5 72B Instruct · 73B · 16.2%Phi 4 · 15B · 11.6%Qwen2.5 7B Instruct · 8B · 7.8%Qwen2 VL 7B Instruct · 8B · 3.7%Llama 3.2 1B Instruct · 1B · 6.6%Llama 3.2 1B InstructLlama 3.2 3B Instruct · 3B · 10.1%Llama 3.2 3B InstructLlama 3.1 8B Instruct · 8B · 15.1%Llama 3.2 11B Vision Instruct · 11B · 16.8%Llama 3.2 11B Vision …Mistral Nemo Instruct 2407 · 12B · 17.6%Mistral Nemo Instruct…DeepSeek R1 Distill Qwen 32B · 33B · 19.5%DeepSeek R1 Distill Q…Llama 3.1 70B Instruct · 71B · 27.9%Llama 3.1 70B InstructDeepSeek R1 · 684B · 34.9%DeepSeek R1
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Llama 3.2 1B Instruct, 1B, score 6.6% — on the efficiency frontier (best score at its size or smaller).
  • Llama 3.2 3B Instruct, 3B, score 10.1% — on the efficiency frontier (best score at its size or smaller).
  • Llama 3.1 8B Instruct, 8B, score 15.1% — on the efficiency frontier (best score at its size or smaller).
  • Llama 3.2 11B Vision Instruct, 11B, score 16.8% — on the efficiency frontier (best score at its size or smaller).
  • Mistral Nemo Instruct 2407, 12B, score 17.6% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek R1 Distill Qwen 32B, 33B, score 19.5% — on the efficiency frontier (best score at its size or smaller).
  • Llama 3.1 70B Instruct, 71B, score 27.9% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek R1, 684B, score 34.9% — on the efficiency frontier (best score at its size or smaller).

BALROG: frequently asked questions

What is the best open LLM on BALROG?
DeepSeek R1 is the top open model on BALROG, scoring 34.9%. Among all models tested — including proprietary ones — it ranks #14. The top model overall is GPT 6 Astra Max (OpenAI) at 68.3%.
What's the best BALROG model you can run on a 24 GB GPU?
DeepSeek R1 Distill Qwen 32B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 18 GB), scoring 19.5% on BALROG.
What's the best BALROG model you can run on a 12 GB GPU?
Mistral Nemo Instruct 2407 is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 7 GB), scoring 17.6% on BALROG.
Can open models match proprietary models on BALROG?
Not quite on BALROG: the strongest proprietary model (GPT 6 Astra Max) scores 68.3%, ahead of the best open model (DeepSeek R1) at 34.9% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.