Reasoning

Fiction.LiveBench Leaderboard

Fiction.LiveBench is a long-context comprehension test: models answer questions that require tracking characters, events and relationships across a long story. llmrun shows the 16k-token score, the column Epoch AI reports as its headline.

Source: epoch24 open models ranked+34 proprietaryData through Jan 2026

All models ranked on Fiction.LiveBench

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1GPT 5 (Aug 07, 2025, medium) · proprietary
97.2%
2O3 Pro (Jun 10, 2025, medium) · proprietary
97.2%
3Grok 4 (Jul 09) · proprietary
94.4%
4Grok 4 Fast · proprietary
94.4%
5Gemini 2.5 Pro Preview (Jun 05) · proprietary
91.7%
6O3 (Apr 16, 2025, medium) · proprietary
88.9%
7Kimi K2.5 · 1026.9B
86.1%
8Claude 3.7 Sonnet (Feb 19, 2025, 8K) · proprietary
83.3%
9DeepSeek V3.2 Exp · 685.4B
83.3%
10O1 (Dec 17, 2024, medium) · proprietary
83.3%
11QwQ 32B · 32.8B
83.3%
12Gemini 2.5 Flash Preview (May 20) · proprietary
77.8%
13O4 Mini (Apr 16, 2025, medium) · proprietary
77.8%
14DeepSeek R1 0528 · 684.5B
75.0%
15Qwen3 235B A22B Thinking 2507 · 235.1B
75.0%
16Qwen3 32B · 32.8B
74.2%
17DeepSeek R1 · 684.5B
69.4%
18GPT 5 Mini (Aug 07, 2025, medium) · proprietary
69.4%
19MiniMax M1 80k · 456.1B
69.4%
20Qwen3 235B A22B · 235.1B
67.7%
21ChatGPT 4o (Jan 29, 2025) · proprietary
66.7%
22Gemini 2.5 Pro Exp (Mar 25) · proprietary
66.7%
23Gemini 2.5 Pro Preview (Mar 25) · proprietary
66.7%
24Gemini 2.5 Pro Preview (May 06) · proprietary
66.7%
25Grok 3 Mini Beta (medium) · proprietary
66.7%
26Kimi K2 Instruct 0905 · 1026.5B
66.7%
27Qwen Max (Jan 25, 2025) · proprietary
66.7%
28Qwen3 Max (Sep 23, 2025) · proprietary
66.7%
29GPT 4.1 (Apr 14, 2025) · proprietary
63.9%
30GPT 4.5 Preview (Feb 27, 2025) · proprietary
63.9%
31Qwen3 14B · 14.8B
62.5%
32Qwen3 8B · 8.2B
62.1%
33Claude Opus 4 (May 14, 2025) · proprietary
61.1%
34Gemini 2.0 Flash 001 · proprietary
61.1%
35Kimi K2 Instruct · 1026.4B
61.1%
36GLM 4.5 · 358.3B
58.3%
37Grok 3 Beta · proprietary
58.3%
38Qwen3 Next 80B A3B Instruct · 81.3B
55.6%
39Parasail Qwen3 235B A22B Instruct 2507 · proprietary
52.9%
40DeepSeek V3.1 · 684.5B
52.8%
41Gemini 2.0 Flash Thinking Exp (Jan 21) · proprietary
52.8%
42Claude 3.7 Sonnet (Feb 19, 2025) · proprietary
50.0%
43DeepSeek v3 0324 · 684.5B
50.0%
44O3 Mini (Jan 31, 2025, medium) · proprietary
50.0%
45Gemini 2.5 Flash Lite Preview Thinking (Jun 17) · proprietary
47.2%
46Gemini 2.5 Flash Preview (Apr 17) · proprietary
47.2%
47Claude Sonnet 4 (May 14, 2025) · proprietary
46.9%
48Llama 4 Maverick 17B 128E Instruct · 401.6B
46.2%
49GPT 4.1 Mini (Apr 14, 2025) · proprietary
44.4%
50GPT 5 Nano (Aug 07, 2025, medium) · proprietary
44.4%
51GPT OSS 120B · 116.8B
44.4%
52Gemini 2.0 Pro Exp (Feb 05) · proprietary
41.7%
53Qwen3 30B A3B · 30.5B
40.6%
54Llama 4 Scout 17B 16E Instruct · 108.6B
36.0%
55Gemma 3 27B IT · 27.4B
33.3%
56Llama 3.3 70B Instruct · 70.6B
33.3%
57GPT 4.1 Nano (Apr 14, 2025) · proprietary
25.0%
58NVIDIA Nemotron Nano 9B v2 · 8.9B
25.0%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

10B100B1Tmodel size (log scale) →86.1%25.0%DeepSeek V3.2 Exp · 685B · 83.3%DeepSeek R1 0528 · 685B · 75.0%Qwen3 235B A22B Thinking 2507 · 235B · 75.0%DeepSeek R1 · 684B · 69.4%MiniMax M1 80k · 456B · 69.4%Qwen3 235B A22B · 235B · 67.7%Kimi K2 Instruct 0905 · 1T · 66.7%Kimi K2 Instruct · 1T · 61.1%GLM 4.5 · 358B · 58.3%Qwen3 Next 80B A3B Instruct · 81B · 55.6%DeepSeek V3.1 · 685B · 52.8%DeepSeek v3 0324 · 685B · 50.0%Llama 4 Maverick 17B 128E Instruct · 402B · 46.2%GPT OSS 120B · 117B · 44.4%Qwen3 30B A3B · 31B · 40.6%Llama 4 Scout 17B 16E Instruct · 109B · 36.0%Gemma 3 27B IT · 27B · 33.3%Llama 3.3 70B Instruct · 71B · 33.3%NVIDIA Nemotron Nano 9B v2 · 9B · 25.0%Qwen3 8B · 8B · 62.1%Qwen3 8BQwen3 14B · 15B · 62.5%Qwen3 14BQwen3 32B · 33B · 74.2%Qwen3 32BQwQ 32B · 33B · 83.3%QwQ 32BKimi K2.5 · 1T · 86.1%Kimi K2.5
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Qwen3 8B, 8B, score 62.1% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3 14B, 15B, score 62.5% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3 32B, 33B, score 74.2% — on the efficiency frontier (best score at its size or smaller).
  • QwQ 32B, 33B, score 83.3% — on the efficiency frontier (best score at its size or smaller).
  • Kimi K2.5, 1T, score 86.1% — on the efficiency frontier (best score at its size or smaller).

Fiction.LiveBench: frequently asked questions

What is the best open LLM on Fiction.LiveBench?
Kimi K2.5 is the top open model on Fiction.LiveBench, scoring 86.1%. Among all models tested — including proprietary ones — it ranks #7. The top model overall is O3 Pro (Jun 10, 2025, medium) (OpenAI) at 97.2%.
What's the best Fiction.LiveBench model you can run on a 24 GB GPU?
QwQ 32B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 18 GB), scoring 83.3% on Fiction.LiveBench.
What's the best Fiction.LiveBench model you can run on a 12 GB GPU?
Qwen3 14B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 8 GB), scoring 62.5% on Fiction.LiveBench.
Can open models match proprietary models on Fiction.LiveBench?
Not quite on Fiction.LiveBench: the strongest proprietary model (O3 Pro (Jun 10, 2025, medium)) scores 97.2%, ahead of the best open model (Kimi K2.5) at 86.1% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.