Reasoning
ForecastBench Leaderboard
ForecastBench, from the Forecasting Research Institute, asks models to predict real-world future events, with questions that only resolve after the model's training cutoff so the answers can't have been memorised. It is scored as a difficulty-adjusted Brier index, where higher is better.
Source: epoch26 open models ranked+56 proprietaryData through Jul 2026
All models ranked on ForecastBench
Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.
| # | Model | Score |
|---|---|---|
| 1 | O3 (Apr 16, 2025, unspecified) · proprietary | 62.5 |
| 2 | Claude Opus 4.1 (Aug 05, 2025, unspecified) · proprietary | 62.0 |
| 3 | Claude Sonnet 4.6 (16K) · proprietary | 62.0 |
| 4 | Claude Sonnet 4.5 (Sep 29, 2025, unspecified) · proprietary | 61.9 |
| 5 | Claude 3.7 Sonnet (Feb 19, 2025, unspecified) · proprietary | 61.8 |
| 6 | O4 Mini (Apr 16, 2025, unspecified) · proprietary | 61.8 |
| 7 | GPT 4.5 Preview (Feb 27, 2025) · proprietary | 61.7 |
| 8 | GPT 4.1 (Apr 14, 2025) · proprietary | 61.5 |
| 9 | Claude Haiku 4.5 (Oct 01, 2025, unspecified) · proprietary | 61.4 |
| 10 | GPT 5 (Aug 07, 2025, unspecified) · proprietary | 61.4 |
| 11 | Grok 4.20 · proprietary | 61.4 |
| 12 | MiniMax M3 · 427.0B | 61.4 |
| 13 | Gemini 2.5 Pro Preview (Mar 25) · proprietary | 61.3 |
| 14 | Gemini 3 Pro Preview · proprietary | 61.2 |
| 15 | Claude Opus 4 (May 14, 2025, unspecified) · proprietary | 61.1 |
| 16 | Claude Sonnet 5 (16K) · proprietary | 61.1 |
| 17 | Kimi K3 · 2779.9B | 61.1 |
| 18 | GLM 5 · 753.9B | 61.0 |
| 19 | GPT 5 Mini (Aug 07, 2025, unspecified) · proprietary | 61.0 |
| 20 | Grok 4.1 Fast Reasoning · proprietary | 61.0 |
| 21 | Grok 4 (Jul 09) · proprietary | 60.9 |
| 22 | Claude 3.5 Sonnet (Oct 22, 2024) · proprietary | 60.7 |
| 23 | Claude Opus 4.5 (Nov 01, 2025, unspecified) · proprietary | 60.7 |
| 24 | Grok 4.20 0309 Reasoning · proprietary | 60.7 |
| 25 | Gemini 2.5 Flash Preview (Apr 17) · proprietary | 60.6 |
| 26 | GPT 5.5 (unspecified) · proprietary | 60.6 |
| 27 | Grok 4 Fast · proprietary | 60.5 |
| 28 | Gemini 2.5 Pro · proprietary | 60.4 |
| 29 | Claude Opus 4.7 (unspecified) · proprietary | 60.3 |
| 30 | Grok 4.3 (unspecified) · proprietary | 60.3 |
| 31 | Claude Sonnet 4 (May 14, 2025, unspecified) · proprietary | 60.2 |
| 32 | Kimi K2 Instruct · 1026.4B | 60.2 |
| 33 | GPT 5.2 (Dec 11, 2025, unspecified) · proprietary | 60.1 |
| 34 | Claude Opus 4.6 (unspecified) · proprietary | 60.0 |
| 35 | DeepSeek R1 · 684.5B | 60.0 |
| 36 | Claude Opus 4.8 (24K) · proprietary | 59.9 |
| 37 | Llama 3.1 405B Instruct · 405.9B | 59.9 |
| 38 | Kimi K2 Instruct 0905 · 1026.5B | 59.8 |
| 39 | Qwen3 235B A22B · 235.1B | 59.7 |
| 40 | Claude Sonnet 4.6 (unspecified) · proprietary | 59.6 |
| 41 | O3 Mini (Jan 31, 2025, unspecified) · proprietary | 59.6 |
| 42 | GPT 5.4 (Mar 05, 2026, unspecified) · proprietary | 59.5 |
| 43 | Claude 3.5 Sonnet (Jun 20, 2024) · proprietary | 59.4 |
| 44 | GPT 4 Turbo (Apr 09, 2024) · proprietary | 59.4 |
| 45 | GLM 4.5 Air · 110.5B | 59.2 |
| 46 | Claude Opus 4.8 (unspecified) · proprietary | 59.1 |
| 47 | DeepSeek v3 · 684.5B | 59.1 |
| 48 | GPT 5 Nano (Aug 07, 2025, unspecified) · proprietary | 59.1 |
| 49 | Gemini 2.5 Flash · proprietary | 59.0 |
| 50 | Gemini 3.1 Pro Preview · proprietary | 59.0 |
| 51 | Gemini 3.5 Flash (unspecified) · proprietary | 59.0 |
| 52 | Llama 3.3 70B Instruct · 70.6B | 58.6 |
| 53 | Meta Llama 3 8B Instruct · 8.0B | 58.6 |
| 54 | Gemini 3 Flash Preview · proprietary | 58.5 |
| 55 | Claude 3 Opus (Feb 29, 2024) · proprietary | 58.4 |
| 56 | Gemini 1.5 Pro 001 · proprietary | 58.4 |
| 57 | QwQ 32B Preview · 32.8B | 58.3 |
| 58 | Gemini 3.1 Pro Preview (high) · proprietary | 58.1 |
| 59 | GPT 5.1 (Nov 13, 2025, unspecified) · proprietary | 58.1 |
| 60 | DeepSeek V3.1 · 684.5B | 58.0 |
| 61 | GPT 4 (Jun 13) · proprietary | 57.8 |
| 62 | GPT 4o (May 13, 2024) · proprietary | 57.7 |
| 63 | Qwen1.5 110B Chat · 111.2B | 57.7 |
| 64 | Llama 4 Maverick 17B 128E Instruct · 401.6B | 57.5 |
| 65 | Llama 4 Scout 17B 16E Instruct · 108.6B | 57.5 |
| 66 | Qwen2.5 72B Instruct · 72.7B | 57.5 |
| 67 | GPT 5.4 Nano (Mar 17, 2026, unspecified) · proprietary | 57.3 |
| 68 | Gemini 2.0 Flash Lite 001 · proprietary | 57.1 |
| 69 | Meta Llama 3 70B · 70.6B | 57.1 |
| 70 | Mistral Large Instruct 2407 · 122.6B | 57.1 |
| 71 | GPT 5.4 Mini (Mar 17, 2026, unspecified) · proprietary | 57.0 |
| 72 | Mistral Large Instruct 2411 · 122.6B | 56.9 |
| 73 | Mixtral 8x22B Instruct v0.1 · 140.6B | 56.3 |
| 74 | Mixtral 8x7B Instruct v0.1 · 46.7B | 56.3 |
| 75 | DeepSeek V4 Pro · 1598.8B | 56.1 |
| 76 | Gemini 3.1 Flash Lite · proprietary | 54.4 |
| 77 | Claude 2.1 · proprietary | 54.2 |
| 78 | Gemini 1.5 Flash 001 · proprietary | 53.9 |
| 79 | Claude 3 Haiku (Mar 07, 2024) · proprietary | 53.2 |
| 80 | Meta Llama 3 8B · 8.0B | 52.9 |
| 81 | Llama 2 70B Chat HF · 69.0B | 51.4 |
| 82 | GPT 3.5 Turbo (Jan 25) · proprietary | 50.4 |
Score vs model size
Which models give the most quality for their size — the ones worth running locally.
- Meta Llama 3 8B Instruct, 8B, score 58.6 — on the efficiency frontier (best score at its size or smaller).
- GLM 4.5 Air, 110B, score 59.2 — on the efficiency frontier (best score at its size or smaller).
- Qwen3 235B A22B, 235B, score 59.7 — on the efficiency frontier (best score at its size or smaller).
- Llama 3.1 405B Instruct, 406B, score 59.9 — on the efficiency frontier (best score at its size or smaller).
- MiniMax M3, 427B, score 61.4 — on the efficiency frontier (best score at its size or smaller).
ForecastBench: frequently asked questions
- What is the best open LLM on ForecastBench?
- MiniMax M3 is the top open model on ForecastBench, scoring 61.4. Among all models tested — including proprietary ones — it ranks #9. The top model overall is O3 (Apr 16, 2025, unspecified) (OpenAI) at 62.5.
- What's the best ForecastBench model you can run on a 24 GB GPU?
- Meta Llama 3 8B Instruct is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 4 GB), scoring 58.6 on ForecastBench.
- What's the best ForecastBench model you can run on a 12 GB GPU?
- Meta Llama 3 8B Instruct is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 4 GB), scoring 58.6 on ForecastBench.
- Can open models match proprietary models on ForecastBench?
- Not quite on ForecastBench: the strongest proprietary model (O3 (Apr 16, 2025, unspecified)) scores 62.5, ahead of the best open model (MiniMax M3) at 61.4 — but you can run the open one yourself.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.