Reasoning

Vending-Bench 2 Leaderboard

Vending-Bench 2, from Andon Labs, is a long-horizon agent benchmark in which the model runs a simulated vending-machine business for a year, handling suppliers, pricing and cash flow. The score is the final money balance in US dollars.

Source: epoch17 open models ranked+46 proprietaryData through Sep 2026

All models ranked on Vending-Bench 2

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1GPT 6 Astra (unspecified) · proprietary
15514.7
2GPT 6 Sol (unspecified) · proprietary
14427.8
3Claude Opus 5 (unspecified) · proprietary
11181.9
4Claude Opus 4.7 (unspecified) · proprietary
10936.8
5GPT 5.6 Sol (unspecified) · proprietary
9619.4
6Grok 4.6 (unspecified) · proprietary
9047.0
7GLM 5.2 · 753.3B
8313.8
8GLM 5.3 · 753.3B
8163.6
9Claude Opus 4.6 (unspecified) · proprietary
8017.6
10GPT 5.5 (unspecified) · proprietary
7523.8
11GPT 5.6 Terra (unspecified) · proprietary
7343.2
12Claude Sonnet 4.6 (unspecified) · proprietary
7204.1
13Muse Spark 1.1 · proprietary
6520.5
14Claude Sonnet 5 (unspecified) · proprietary
6377.7
15Kimi K2.6 · 1026.9B
6204.6
16GPT 5.4 (Mar 05, 2026, unspecified) · proprietary
6144.2
17GPT 5.3 Codex · proprietary
5940.1
18Claude Opus 4.8 (unspecified) · proprietary
5787.4
19Claude Fable 5 (high) · proprietary
5680.3
20GLM 5.1 · 753.9B
5634.4
21Gemini 3 Pro Preview · proprietary
5478.2
22Claude Fable 5.1 (unspecified) · proprietary
5421.6
23Gemini 3.5 Flash (unspecified) · proprietary
5396.4
24Kimi K3 · 2779.9B
5165.0
25Qwen3.6 Plus · proprietary
5114.9
26Gemini 3.8 Flash (unspecified) · proprietary
5093.8
27Kimi K2.7 Code · 1026.9B
5082.9
28Claude Fable 5 (low) · proprietary
5018.5
29Claude Opus 4.5 (Nov 01, 2025, unspecified) · proprietary
4967.1
30Claude Fable 5 Max · proprietary
4966.6
31Grok 4.20 · proprietary
4662.8
32Claude Fable 5 · proprietary
4529.9
33GLM 5 · 753.9B
4432.1
34Claude Fable 5 (medium) · proprietary
4339.8
35Qwen3.6 Max Preview · proprietary
4254.2
36GPT 5.6 Luna (unspecified) · proprietary
4094.7
37Grok 4.5 (unspecified) · proprietary
3887.4
38Claude Sonnet 4.5 (Sep 29, 2025, unspecified) · proprietary
3838.7
39Gemini 3.1 Pro Preview Customtools · proprietary
3774.3
40Gemini 3 Flash Preview · proprietary
3634.7
41GPT 5.2 (Dec 11, 2025, unspecified) · proprietary
3591.3
42DeepSeek V4 Pro · 1598.8B
3284.5
43Claude Opus 4.8 Max · proprietary
2992.3
44GLM 4.7 · 358.3B
2376.8
45MiniMax M3 · 427.0B
2157.8
46GPT 5.1 (Nov 13, 2025, unspecified) · proprietary
1473.4
47Kimi K2.5 · 1026.9B
1198.5
48Grok 4.1 Fast Reasoning · proprietary
1106.6
49DeepSeek V3.2 Exp · 685.4B
1034.0
50Gemini 3.1 Pro Preview · proprietary
911.2
51Gemini 2.5 Pro · proprietary
573.6
52Gemini 2.5 Flash · proprietary
548.8
53Qwen3.5 Flash · proprietary
462.7
54Claude Haiku 4.5 (Oct 01, 2025, unspecified) · proprietary
458.9
55Qwen3.5 27B · 27.8B
202.0
56MiniMax M2 · 228.7B
160.6
57Qwen3 Max (Sep 23, 2025) · proprietary
71.6
58Grok 4.3 (unspecified) · proprietary
35.3
59Qwen3.5 Plus · proprietary
0.5
60Qwen3 235B A22B Thinking 2507 · 235.1B
-11.3
61GPT OSS 120B · 116.8B
-21.5
62MiniMax M2.5 · 228.7B
-23.2
63GPT 5 Mini (Aug 07, 2025, unspecified) · proprietary
-31.2

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

100B1Tmodel size (log scale) →8313.8-23.2GLM 5.3 · 753B · 8163.6Kimi K2.6 · 1T · 6204.6GLM 5.1 · 754B · 5634.4Kimi K3 · 2.8T · 5165.0Kimi K2.7 Code · 1T · 5082.9GLM 5 · 754B · 4432.1DeepSeek V4 Pro · 1.6T · 3284.5MiniMax M3 · 427B · 2157.8Kimi K2.5 · 1T · 1198.5DeepSeek V3.2 Exp · 685B · 1034.0MiniMax M2 · 229B · 160.6Qwen3 235B A22B Thinking 2507 · 235B · -11.3GPT OSS 120B · 117B · -21.5MiniMax M2.5 · 229B · -23.2Qwen3.5 27B · 28B · 202.0Qwen3.5 27BGLM 4.7 · 358B · 2376.8GLM 4.7GLM 5.2 · 753B · 8313.8GLM 5.2
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Qwen3.5 27B, 28B, score 202.0 — on the efficiency frontier (best score at its size or smaller).
  • GLM 4.7, 358B, score 2376.8 — on the efficiency frontier (best score at its size or smaller).
  • GLM 5.2, 753B, score 8313.8 — on the efficiency frontier (best score at its size or smaller).

Vending-Bench 2: frequently asked questions

What is the best open LLM on Vending-Bench 2?
GLM 5.2 is the top open model on Vending-Bench 2, scoring 8313.8. Among all models tested — including proprietary ones — it ranks #7. The top model overall is GPT 6 Astra (unspecified) (OpenAI) at 15514.7.
What's the best Vending-Bench 2 model you can run on a 24 GB GPU?
Qwen3.5 27B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 15 GB), scoring 202.0 on Vending-Bench 2.
Can open models match proprietary models on Vending-Bench 2?
Not quite on Vending-Bench 2: the strongest proprietary model (GPT 6 Astra (unspecified)) scores 15514.7, ahead of the best open model (GLM 5.2) at 8313.8 — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.