Reasoning
Fiction.LiveBench Leaderboard
Fiction.LiveBench is a long-context comprehension test: models answer questions that require tracking characters, events and relationships across a long story. llmrun shows the 16k-token score, the column Epoch AI reports as its headline.
Source: epoch24 open models ranked+34 proprietaryData through Jan 2026
All models ranked on Fiction.LiveBench
Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.
| # | Model | Score |
|---|---|---|
| 1 | GPT 5 (Aug 07, 2025, medium) · proprietary | 97.2% |
| 2 | O3 Pro (Jun 10, 2025, medium) · proprietary | 97.2% |
| 3 | Grok 4 (Jul 09) · proprietary | 94.4% |
| 4 | Grok 4 Fast · proprietary | 94.4% |
| 5 | Gemini 2.5 Pro Preview (Jun 05) · proprietary | 91.7% |
| 6 | O3 (Apr 16, 2025, medium) · proprietary | 88.9% |
| 7 | Kimi K2.5 · 1026.9B | 86.1% |
| 8 | Claude 3.7 Sonnet (Feb 19, 2025, 8K) · proprietary | 83.3% |
| 9 | DeepSeek V3.2 Exp · 685.4B | 83.3% |
| 10 | O1 (Dec 17, 2024, medium) · proprietary | 83.3% |
| 11 | QwQ 32B · 32.8B | 83.3% |
| 12 | Gemini 2.5 Flash Preview (May 20) · proprietary | 77.8% |
| 13 | O4 Mini (Apr 16, 2025, medium) · proprietary | 77.8% |
| 14 | DeepSeek R1 0528 · 684.5B | 75.0% |
| 15 | Qwen3 235B A22B Thinking 2507 · 235.1B | 75.0% |
| 16 | Qwen3 32B · 32.8B | 74.2% |
| 17 | DeepSeek R1 · 684.5B | 69.4% |
| 18 | GPT 5 Mini (Aug 07, 2025, medium) · proprietary | 69.4% |
| 19 | MiniMax M1 80k · 456.1B | 69.4% |
| 20 | Qwen3 235B A22B · 235.1B | 67.7% |
| 21 | ChatGPT 4o (Jan 29, 2025) · proprietary | 66.7% |
| 22 | Gemini 2.5 Pro Exp (Mar 25) · proprietary | 66.7% |
| 23 | Gemini 2.5 Pro Preview (Mar 25) · proprietary | 66.7% |
| 24 | Gemini 2.5 Pro Preview (May 06) · proprietary | 66.7% |
| 25 | Grok 3 Mini Beta (medium) · proprietary | 66.7% |
| 26 | Kimi K2 Instruct 0905 · 1026.5B | 66.7% |
| 27 | Qwen Max (Jan 25, 2025) · proprietary | 66.7% |
| 28 | Qwen3 Max (Sep 23, 2025) · proprietary | 66.7% |
| 29 | GPT 4.1 (Apr 14, 2025) · proprietary | 63.9% |
| 30 | GPT 4.5 Preview (Feb 27, 2025) · proprietary | 63.9% |
| 31 | Qwen3 14B · 14.8B | 62.5% |
| 32 | Qwen3 8B · 8.2B | 62.1% |
| 33 | Claude Opus 4 (May 14, 2025) · proprietary | 61.1% |
| 34 | Gemini 2.0 Flash 001 · proprietary | 61.1% |
| 35 | Kimi K2 Instruct · 1026.4B | 61.1% |
| 36 | GLM 4.5 · 358.3B | 58.3% |
| 37 | Grok 3 Beta · proprietary | 58.3% |
| 38 | Qwen3 Next 80B A3B Instruct · 81.3B | 55.6% |
| 39 | Parasail Qwen3 235B A22B Instruct 2507 · proprietary | 52.9% |
| 40 | DeepSeek V3.1 · 684.5B | 52.8% |
| 41 | Gemini 2.0 Flash Thinking Exp (Jan 21) · proprietary | 52.8% |
| 42 | Claude 3.7 Sonnet (Feb 19, 2025) · proprietary | 50.0% |
| 43 | DeepSeek v3 0324 · 684.5B | 50.0% |
| 44 | O3 Mini (Jan 31, 2025, medium) · proprietary | 50.0% |
| 45 | Gemini 2.5 Flash Lite Preview Thinking (Jun 17) · proprietary | 47.2% |
| 46 | Gemini 2.5 Flash Preview (Apr 17) · proprietary | 47.2% |
| 47 | Claude Sonnet 4 (May 14, 2025) · proprietary | 46.9% |
| 48 | Llama 4 Maverick 17B 128E Instruct · 401.6B | 46.2% |
| 49 | GPT 4.1 Mini (Apr 14, 2025) · proprietary | 44.4% |
| 50 | GPT 5 Nano (Aug 07, 2025, medium) · proprietary | 44.4% |
| 51 | GPT OSS 120B · 116.8B | 44.4% |
| 52 | Gemini 2.0 Pro Exp (Feb 05) · proprietary | 41.7% |
| 53 | Qwen3 30B A3B · 30.5B | 40.6% |
| 54 | Llama 4 Scout 17B 16E Instruct · 108.6B | 36.0% |
| 55 | Gemma 3 27B IT · 27.4B | 33.3% |
| 56 | Llama 3.3 70B Instruct · 70.6B | 33.3% |
| 57 | GPT 4.1 Nano (Apr 14, 2025) · proprietary | 25.0% |
| 58 | NVIDIA Nemotron Nano 9B v2 · 8.9B | 25.0% |
Score vs model size
Which models give the most quality for their size — the ones worth running locally.
- Qwen3 8B, 8B, score 62.1% — on the efficiency frontier (best score at its size or smaller).
- Qwen3 14B, 15B, score 62.5% — on the efficiency frontier (best score at its size or smaller).
- Qwen3 32B, 33B, score 74.2% — on the efficiency frontier (best score at its size or smaller).
- QwQ 32B, 33B, score 83.3% — on the efficiency frontier (best score at its size or smaller).
- Kimi K2.5, 1T, score 86.1% — on the efficiency frontier (best score at its size or smaller).
Fiction.LiveBench: frequently asked questions
- What is the best open LLM on Fiction.LiveBench?
- Kimi K2.5 is the top open model on Fiction.LiveBench, scoring 86.1%. Among all models tested — including proprietary ones — it ranks #7. The top model overall is O3 Pro (Jun 10, 2025, medium) (OpenAI) at 97.2%.
- What's the best Fiction.LiveBench model you can run on a 24 GB GPU?
- QwQ 32B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 18 GB), scoring 83.3% on Fiction.LiveBench.
- What's the best Fiction.LiveBench model you can run on a 12 GB GPU?
- Qwen3 14B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 8 GB), scoring 62.5% on Fiction.LiveBench.
- Can open models match proprietary models on Fiction.LiveBench?
- Not quite on Fiction.LiveBench: the strongest proprietary model (O3 Pro (Jun 10, 2025, medium)) scores 97.2%, ahead of the best open model (Kimi K2.5) at 86.1% — but you can run the open one yourself.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.