Reasoning
BALROG Leaderboard
BALROG (Benchmarking Agentic LLM and VLM Reasoning On Games) tests long-horizon agentic reasoning across six game environments — BabyAI, Crafter, TextWorld, Baba Is AI, MiniHack and NetHack — scoring the average progress a model makes toward completing each game. It was built by UCL's DARK lab and collaborators.
Source: epoch14 open models ranked+26 proprietaryData through Sep 2026
All models ranked on BALROG
Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.
| # | Model | Score |
|---|---|---|
| 1 | GPT 6 Astra Max · proprietary | 68.3% |
| 2 | Claude Opus 5 Max · proprietary | 63.4% |
| 3 | GPT 5.6 Sol Max · proprietary | 60.0% |
| 4 | Gemini 3 Pro Preview · proprietary | 58.1% |
| 5 | Gemini 3.1 Pro Preview · proprietary | 57.0% |
| 6 | GPT 5.6 Terra Max · proprietary | 53.2% |
| 7 | Gemini 3 Flash Preview · proprietary | 48.1% |
| 8 | GPT 5.6 Luna Max · proprietary | 45.6% |
| 9 | Grok 4 (Jul 09) · proprietary | 43.6% |
| 10 | Claude Opus 4.5 (Nov 01, 2025) · proprietary | 43.5% |
| 11 | Gemini 2.5 Pro Exp (Mar 25) · proprietary | 43.3% |
| 12 | Claude Opus 4.5 (Nov 01, 2025, 64K) · proprietary | 43.0% |
| 13 | Claude Opus 4.5 (Nov 01, 2025, unspecified) · proprietary | 43.0% |
| 14 | DeepSeek R1 · 684.5B | 34.9% |
| 15 | Gemini 2.5 Flash · proprietary | 33.5% |
| 16 | GPT 5 (Aug 07, 2025, minimal) · proprietary | 32.8% |
| 17 | Claude 3.5 Sonnet (Oct 22, 2024) · proprietary | 32.6% |
| 18 | GPT 4o (May 13, 2024) · proprietary | 32.3% |
| 19 | Claude Haiku 4.5 (Oct 01, 2025, 1K) · proprietary | 31.2% |
| 20 | Claude Haiku 4.5 (Oct 01, 2025) · proprietary | 31.2% |
| 21 | Grok 3 Beta · proprietary | 29.5% |
| 22 | Reka Flash 3 · proprietary | 29.2% |
| 23 | Llama 3.1 70B Instruct · 70.6B | 27.9% |
| 24 | Llama 3.2 90B Vision Instruct · 88.6B | 27.3% |
| 25 | Llama 3.3 70B Instruct · 70.6B | 23.0% |
| 26 | Gemini 1.5 Pro 002 · proprietary | 21.0% |
| 27 | DeepSeek R1 Distill Qwen 32B · 32.8B | 19.5% |
| 28 | Claude 3.5 Haiku (Oct 22, 2024) · proprietary | 19.3% |
| 29 | Mistral Nemo Instruct 2407 · 12.2B | 17.6% |
| 30 | GPT 4o Mini (Jul 18, 2024) · proprietary | 17.4% |
| 31 | Llama 3.2 11B Vision Instruct · 10.7B | 16.8% |
| 32 | Qwen2.5 72B Instruct · 72.7B | 16.2% |
| 33 | Llama 3.1 8B Instruct · 8.0B | 15.1% |
| 34 | Gemini 1.5 Flash 002 · proprietary | 14.6% |
| 35 | Qwen2 VL 72B Instruct · proprietary | 12.8% |
| 36 | Phi 4 · 14.7B | 11.6% |
| 37 | Llama 3.2 3B Instruct · 3.2B | 10.1% |
| 38 | Qwen2.5 7B Instruct · 7.6B | 7.8% |
| 39 | Llama 3.2 1B Instruct · 1.2B | 6.6% |
| 40 | Qwen2 VL 7B Instruct · 8.3B | 3.7% |
Score vs model size
Which models give the most quality for their size — the ones worth running locally.
- Llama 3.2 1B Instruct, 1B, score 6.6% — on the efficiency frontier (best score at its size or smaller).
- Llama 3.2 3B Instruct, 3B, score 10.1% — on the efficiency frontier (best score at its size or smaller).
- Llama 3.1 8B Instruct, 8B, score 15.1% — on the efficiency frontier (best score at its size or smaller).
- Llama 3.2 11B Vision Instruct, 11B, score 16.8% — on the efficiency frontier (best score at its size or smaller).
- Mistral Nemo Instruct 2407, 12B, score 17.6% — on the efficiency frontier (best score at its size or smaller).
- DeepSeek R1 Distill Qwen 32B, 33B, score 19.5% — on the efficiency frontier (best score at its size or smaller).
- Llama 3.1 70B Instruct, 71B, score 27.9% — on the efficiency frontier (best score at its size or smaller).
- DeepSeek R1, 684B, score 34.9% — on the efficiency frontier (best score at its size or smaller).
BALROG: frequently asked questions
- What is the best open LLM on BALROG?
- DeepSeek R1 is the top open model on BALROG, scoring 34.9%. Among all models tested — including proprietary ones — it ranks #14. The top model overall is GPT 6 Astra Max (OpenAI) at 68.3%.
- What's the best BALROG model you can run on a 24 GB GPU?
- DeepSeek R1 Distill Qwen 32B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 18 GB), scoring 19.5% on BALROG.
- What's the best BALROG model you can run on a 12 GB GPU?
- Mistral Nemo Instruct 2407 is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 7 GB), scoring 17.6% on BALROG.
- Can open models match proprietary models on BALROG?
- Not quite on BALROG: the strongest proprietary model (GPT 6 Astra Max) scores 68.3%, ahead of the best open model (DeepSeek R1) at 34.9% — but you can run the open one yourself.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.