Reasoning
Fiction.LiveBench Leaderboard
Fiction.LiveBench is a long-context comprehension test: models answer questions that require tracking characters, events and relationships across a long story. llmrun shows the 16k-token score, the column Epoch AI reports as its headline.
Source: epoch24 open models ranked+34 proprietaryData through Jan 2026
Open models ranked on Fiction.LiveBench
# shows rank among open models / rank overall (including proprietary).
| # | Model | Score |
|---|---|---|
| 1 / 7 | Kimi K2.5 · 1026.9B | 86.1% |
| 2 / 9 | DeepSeek V3.2 Exp · 685.4B | 83.3% |
| 3 / 11 | QwQ 32B · 32.8B | 83.3% |
| 4 / 14 | DeepSeek R1 0528 · 684.5B | 75.0% |
| 5 / 15 | Qwen3 235B A22B Thinking 2507 · 235.1B | 75.0% |
| 6 / 16 | Qwen3 32B · 32.8B | 74.2% |
| 7 / 17 | DeepSeek R1 · 684.5B | 69.4% |
| 8 / 19 | MiniMax M1 80k · 456.1B | 69.4% |
| 9 / 20 | Qwen3 235B A22B · 235.1B | 67.7% |
| 10 / 26 | Kimi K2 Instruct 0905 · 1026.5B | 66.7% |
| 11 / 31 | Qwen3 14B · 14.8B | 62.5% |
| 12 / 32 | Qwen3 8B · 8.2B | 62.1% |
| 13 / 35 | Kimi K2 Instruct · 1026.4B | 61.1% |
| 14 / 36 | GLM 4.5 · 358.3B | 58.3% |
| 15 / 38 | Qwen3 Next 80B A3B Instruct · 81.3B | 55.6% |
| 16 / 40 | DeepSeek V3.1 · 684.5B | 52.8% |
| 17 / 43 | DeepSeek v3 0324 · 684.5B | 50.0% |
| 18 / 48 | Llama 4 Maverick 17B 128E Instruct · 401.6B | 46.2% |
| 19 / 51 | GPT OSS 120B · 116.8B | 44.4% |
| 20 / 53 | Qwen3 30B A3B · 30.5B | 40.6% |
| 21 / 54 | Llama 4 Scout 17B 16E Instruct · 108.6B | 36.0% |
| 22 / 55 | Gemma 3 27B IT · 27.4B | 33.3% |
| 23 / 56 | Llama 3.3 70B Instruct · 70.6B | 33.3% |
| 24 / 58 | NVIDIA Nemotron Nano 9B v2 · 8.9B | 25.0% |
Score vs model size
Which models give the most quality for their size — the ones worth running locally.
- Qwen3 8B, 8B, score 62.1% — on the efficiency frontier (best score at its size or smaller).
- Qwen3 14B, 15B, score 62.5% — on the efficiency frontier (best score at its size or smaller).
- Qwen3 32B, 33B, score 74.2% — on the efficiency frontier (best score at its size or smaller).
- QwQ 32B, 33B, score 83.3% — on the efficiency frontier (best score at its size or smaller).
- Kimi K2.5, 1T, score 86.1% — on the efficiency frontier (best score at its size or smaller).
Fiction.LiveBench: frequently asked questions
- What is the best open LLM on Fiction.LiveBench?
- Kimi K2.5 is the top open model on Fiction.LiveBench, scoring 86.1%. Among all models tested — including proprietary ones — it ranks #7. The top model overall is O3 Pro (Jun 10, 2025, medium) (OpenAI) at 97.2%.
- What's the best Fiction.LiveBench model you can run on a 24 GB GPU?
- QwQ 32B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 18 GB), scoring 83.3% on Fiction.LiveBench.
- What's the best Fiction.LiveBench model you can run on a 12 GB GPU?
- Qwen3 14B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 8 GB), scoring 62.5% on Fiction.LiveBench.
- Can open models match proprietary models on Fiction.LiveBench?
- Not quite on Fiction.LiveBench: the strongest proprietary model (O3 Pro (Jun 10, 2025, medium)) scores 97.2%, ahead of the best open model (Kimi K2.5) at 86.1% — but you can run the open one yourself.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.