Reasoning

Fiction.LiveBench Leaderboard

Fiction.LiveBench is a long-context comprehension test: models answer questions that require tracking characters, events and relationships across a long story. llmrun shows the 16k-token score, the column Epoch AI reports as its headline.

Source: epoch24 open models ranked+34 proprietaryData through Jan 2026

Open models ranked on Fiction.LiveBench

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 7Kimi K2.5 · 1026.9B
86.1%
2 / 9DeepSeek V3.2 Exp · 685.4B
83.3%
3 / 11QwQ 32B · 32.8B
83.3%
4 / 14DeepSeek R1 0528 · 684.5B
75.0%
5 / 15Qwen3 235B A22B Thinking 2507 · 235.1B
75.0%
6 / 16Qwen3 32B · 32.8B
74.2%
7 / 17DeepSeek R1 · 684.5B
69.4%
8 / 19MiniMax M1 80k · 456.1B
69.4%
9 / 20Qwen3 235B A22B · 235.1B
67.7%
10 / 26Kimi K2 Instruct 0905 · 1026.5B
66.7%
11 / 31Qwen3 14B · 14.8B
62.5%
12 / 32Qwen3 8B · 8.2B
62.1%
13 / 35Kimi K2 Instruct · 1026.4B
61.1%
14 / 36GLM 4.5 · 358.3B
58.3%
15 / 38Qwen3 Next 80B A3B Instruct · 81.3B
55.6%
16 / 40DeepSeek V3.1 · 684.5B
52.8%
17 / 43DeepSeek v3 0324 · 684.5B
50.0%
18 / 48Llama 4 Maverick 17B 128E Instruct · 401.6B
46.2%
19 / 51GPT OSS 120B · 116.8B
44.4%
20 / 53Qwen3 30B A3B · 30.5B
40.6%
21 / 54Llama 4 Scout 17B 16E Instruct · 108.6B
36.0%
22 / 55Gemma 3 27B IT · 27.4B
33.3%
23 / 56Llama 3.3 70B Instruct · 70.6B
33.3%
24 / 58NVIDIA Nemotron Nano 9B v2 · 8.9B
25.0%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

10B100B1Tmodel size (log scale) →86.1%25.0%DeepSeek V3.2 Exp · 685B · 83.3%DeepSeek R1 0528 · 685B · 75.0%Qwen3 235B A22B Thinking 2507 · 235B · 75.0%DeepSeek R1 · 684B · 69.4%MiniMax M1 80k · 456B · 69.4%Qwen3 235B A22B · 235B · 67.7%Kimi K2 Instruct 0905 · 1T · 66.7%Kimi K2 Instruct · 1T · 61.1%GLM 4.5 · 358B · 58.3%Qwen3 Next 80B A3B Instruct · 81B · 55.6%DeepSeek V3.1 · 685B · 52.8%DeepSeek v3 0324 · 685B · 50.0%Llama 4 Maverick 17B 128E Instruct · 402B · 46.2%GPT OSS 120B · 117B · 44.4%Qwen3 30B A3B · 31B · 40.6%Llama 4 Scout 17B 16E Instruct · 109B · 36.0%Gemma 3 27B IT · 27B · 33.3%Llama 3.3 70B Instruct · 71B · 33.3%NVIDIA Nemotron Nano 9B v2 · 9B · 25.0%Qwen3 8B · 8B · 62.1%Qwen3 8BQwen3 14B · 15B · 62.5%Qwen3 14BQwen3 32B · 33B · 74.2%Qwen3 32BQwQ 32B · 33B · 83.3%QwQ 32BKimi K2.5 · 1T · 86.1%Kimi K2.5
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Qwen3 8B, 8B, score 62.1% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3 14B, 15B, score 62.5% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3 32B, 33B, score 74.2% — on the efficiency frontier (best score at its size or smaller).
  • QwQ 32B, 33B, score 83.3% — on the efficiency frontier (best score at its size or smaller).
  • Kimi K2.5, 1T, score 86.1% — on the efficiency frontier (best score at its size or smaller).

Fiction.LiveBench: frequently asked questions

What is the best open LLM on Fiction.LiveBench?
Kimi K2.5 is the top open model on Fiction.LiveBench, scoring 86.1%. Among all models tested — including proprietary ones — it ranks #7. The top model overall is O3 Pro (Jun 10, 2025, medium) (OpenAI) at 97.2%.
What's the best Fiction.LiveBench model you can run on a 24 GB GPU?
QwQ 32B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 18 GB), scoring 83.3% on Fiction.LiveBench.
What's the best Fiction.LiveBench model you can run on a 12 GB GPU?
Qwen3 14B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 8 GB), scoring 62.5% on Fiction.LiveBench.
Can open models match proprietary models on Fiction.LiveBench?
Not quite on Fiction.LiveBench: the strongest proprietary model (O3 Pro (Jun 10, 2025, medium)) scores 97.2%, ahead of the best open model (Kimi K2.5) at 86.1% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.