Reasoning

ForecastBench Leaderboard

ForecastBench, from the Forecasting Research Institute, asks models to predict real-world future events, with questions that only resolve after the model's training cutoff so the answers can't have been memorised. It is scored as a difficulty-adjusted Brier index, where higher is better.

Source: epoch26 open models ranked+56 proprietaryData through Jul 2026

All models ranked on ForecastBench

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1O3 (Apr 16, 2025, unspecified) · proprietary
62.5
2Claude Opus 4.1 (Aug 05, 2025, unspecified) · proprietary
62.0
3Claude Sonnet 4.6 (16K) · proprietary
62.0
4Claude Sonnet 4.5 (Sep 29, 2025, unspecified) · proprietary
61.9
5Claude 3.7 Sonnet (Feb 19, 2025, unspecified) · proprietary
61.8
6O4 Mini (Apr 16, 2025, unspecified) · proprietary
61.8
7GPT 4.5 Preview (Feb 27, 2025) · proprietary
61.7
8GPT 4.1 (Apr 14, 2025) · proprietary
61.5
9Claude Haiku 4.5 (Oct 01, 2025, unspecified) · proprietary
61.4
10GPT 5 (Aug 07, 2025, unspecified) · proprietary
61.4
11Grok 4.20 · proprietary
61.4
12MiniMax M3 · 427.0B
61.4
13Gemini 2.5 Pro Preview (Mar 25) · proprietary
61.3
14Gemini 3 Pro Preview · proprietary
61.2
15Claude Opus 4 (May 14, 2025, unspecified) · proprietary
61.1
16Claude Sonnet 5 (16K) · proprietary
61.1
17Kimi K3 · 2779.9B
61.1
18GLM 5 · 753.9B
61.0
19GPT 5 Mini (Aug 07, 2025, unspecified) · proprietary
61.0
20Grok 4.1 Fast Reasoning · proprietary
61.0
21Grok 4 (Jul 09) · proprietary
60.9
22Claude 3.5 Sonnet (Oct 22, 2024) · proprietary
60.7
23Claude Opus 4.5 (Nov 01, 2025, unspecified) · proprietary
60.7
24Grok 4.20 0309 Reasoning · proprietary
60.7
25Gemini 2.5 Flash Preview (Apr 17) · proprietary
60.6
26GPT 5.5 (unspecified) · proprietary
60.6
27Grok 4 Fast · proprietary
60.5
28Gemini 2.5 Pro · proprietary
60.4
29Claude Opus 4.7 (unspecified) · proprietary
60.3
30Grok 4.3 (unspecified) · proprietary
60.3
31Claude Sonnet 4 (May 14, 2025, unspecified) · proprietary
60.2
32Kimi K2 Instruct · 1026.4B
60.2
33GPT 5.2 (Dec 11, 2025, unspecified) · proprietary
60.1
34Claude Opus 4.6 (unspecified) · proprietary
60.0
35DeepSeek R1 · 684.5B
60.0
36Claude Opus 4.8 (24K) · proprietary
59.9
37Llama 3.1 405B Instruct · 405.9B
59.9
38Kimi K2 Instruct 0905 · 1026.5B
59.8
39Qwen3 235B A22B · 235.1B
59.7
40Claude Sonnet 4.6 (unspecified) · proprietary
59.6
41O3 Mini (Jan 31, 2025, unspecified) · proprietary
59.6
42GPT 5.4 (Mar 05, 2026, unspecified) · proprietary
59.5
43Claude 3.5 Sonnet (Jun 20, 2024) · proprietary
59.4
44GPT 4 Turbo (Apr 09, 2024) · proprietary
59.4
45GLM 4.5 Air · 110.5B
59.2
46Claude Opus 4.8 (unspecified) · proprietary
59.1
47DeepSeek v3 · 684.5B
59.1
48GPT 5 Nano (Aug 07, 2025, unspecified) · proprietary
59.1
49Gemini 2.5 Flash · proprietary
59.0
50Gemini 3.1 Pro Preview · proprietary
59.0
51Gemini 3.5 Flash (unspecified) · proprietary
59.0
52Llama 3.3 70B Instruct · 70.6B
58.6
53Meta Llama 3 8B Instruct · 8.0B
58.6
54Gemini 3 Flash Preview · proprietary
58.5
55Claude 3 Opus (Feb 29, 2024) · proprietary
58.4
56Gemini 1.5 Pro 001 · proprietary
58.4
57QwQ 32B Preview · 32.8B
58.3
58Gemini 3.1 Pro Preview (high) · proprietary
58.1
59GPT 5.1 (Nov 13, 2025, unspecified) · proprietary
58.1
60DeepSeek V3.1 · 684.5B
58.0
61GPT 4 (Jun 13) · proprietary
57.8
62GPT 4o (May 13, 2024) · proprietary
57.7
63Qwen1.5 110B Chat · 111.2B
57.7
64Llama 4 Maverick 17B 128E Instruct · 401.6B
57.5
65Llama 4 Scout 17B 16E Instruct · 108.6B
57.5
66Qwen2.5 72B Instruct · 72.7B
57.5
67GPT 5.4 Nano (Mar 17, 2026, unspecified) · proprietary
57.3
68Gemini 2.0 Flash Lite 001 · proprietary
57.1
69Meta Llama 3 70B · 70.6B
57.1
70Mistral Large Instruct 2407 · 122.6B
57.1
71GPT 5.4 Mini (Mar 17, 2026, unspecified) · proprietary
57.0
72Mistral Large Instruct 2411 · 122.6B
56.9
73Mixtral 8x22B Instruct v0.1 · 140.6B
56.3
74Mixtral 8x7B Instruct v0.1 · 46.7B
56.3
75DeepSeek V4 Pro · 1598.8B
56.1
76Gemini 3.1 Flash Lite · proprietary
54.4
77Claude 2.1 · proprietary
54.2
78Gemini 1.5 Flash 001 · proprietary
53.9
79Claude 3 Haiku (Mar 07, 2024) · proprietary
53.2
80Meta Llama 3 8B · 8.0B
52.9
81Llama 2 70B Chat HF · 69.0B
51.4
82GPT 3.5 Turbo (Jan 25) · proprietary
50.4

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

10B100B1Tmodel size (log scale) →61.451.4Kimi K3 · 2.8T · 61.1GLM 5 · 754B · 61.0Kimi K2 Instruct · 1T · 60.2DeepSeek R1 · 684B · 60.0Kimi K2 Instruct 0905 · 1T · 59.8DeepSeek v3 · 685B · 59.1Llama 3.3 70B Instruct · 71B · 58.6QwQ 32B Preview · 33B · 58.3DeepSeek V3.1 · 685B · 58.0Qwen1.5 110B Chat · 111B · 57.7Llama 4 Maverick 17B 128E Instruct · 402B · 57.5Llama 4 Scout 17B 16E Instruct · 109B · 57.5Qwen2.5 72B Instruct · 73B · 57.5Meta Llama 3 70B · 71B · 57.1Mistral Large Instruct 2407 · 123B · 57.1Mistral Large Instruct 2411 · 123B · 56.9Mixtral 8x22B Instruct v0.1 · 141B · 56.3Mixtral 8x7B Instruct v0.1 · 47B · 56.3DeepSeek V4 Pro · 1.6T · 56.1Meta Llama 3 8B · 8B · 52.9Llama 2 70B Chat HF · 69B · 51.4Meta Llama 3 8B Instruct · 8B · 58.6Meta Llama 3 8B Instr…GLM 4.5 Air · 110B · 59.2GLM 4.5 AirQwen3 235B A22B · 235B · 59.7Qwen3 235B A22BLlama 3.1 405B Instruct · 406B · 59.9Llama 3.1 405B Instru…MiniMax M3 · 427B · 61.4MiniMax M3
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Meta Llama 3 8B Instruct, 8B, score 58.6 — on the efficiency frontier (best score at its size or smaller).
  • GLM 4.5 Air, 110B, score 59.2 — on the efficiency frontier (best score at its size or smaller).
  • Qwen3 235B A22B, 235B, score 59.7 — on the efficiency frontier (best score at its size or smaller).
  • Llama 3.1 405B Instruct, 406B, score 59.9 — on the efficiency frontier (best score at its size or smaller).
  • MiniMax M3, 427B, score 61.4 — on the efficiency frontier (best score at its size or smaller).

ForecastBench: frequently asked questions

What is the best open LLM on ForecastBench?
MiniMax M3 is the top open model on ForecastBench, scoring 61.4. Among all models tested — including proprietary ones — it ranks #9. The top model overall is O3 (Apr 16, 2025, unspecified) (OpenAI) at 62.5.
What's the best ForecastBench model you can run on a 24 GB GPU?
Meta Llama 3 8B Instruct is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 4 GB), scoring 58.6 on ForecastBench.
What's the best ForecastBench model you can run on a 12 GB GPU?
Meta Llama 3 8B Instruct is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 4 GB), scoring 58.6 on ForecastBench.
Can open models match proprietary models on ForecastBench?
Not quite on ForecastBench: the strongest proprietary model (O3 (Apr 16, 2025, unspecified)) scores 62.5, ahead of the best open model (MiniMax M3) at 61.4 — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.