Coding

WeirdML Leaderboard

WeirdML asks models to write machine-learning code for small, deliberately unusual datasets, then runs the code and scores the accuracy it actually achieves. Built by Håvard Tveit Ihle, it rewards practical problem-solving over recognizing a familiar task.

Source: epoch46 open models ranked+115 proprietaryData through Sep 2026

All models ranked on WeirdML

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1GPT 6 Astra Promax · proprietary
93.6%
2GPT 6 Astra Max · proprietary
93.3%
3Claude Fable 5.1 Max · proprietary
92.9%
4GPT 6 Astra (high) · proprietary
92.9%
5Claude Fable 5.1 (high) · proprietary
92.3%
6Claude Fable 5 Max · proprietary
91.9%
7Claude Opus 5 Max · proprietary
91.8%
8Claude Opus 5 (high) · proprietary
91.6%
9GPT 5.6 Sol Promax · proprietary
89.4%
10GPT 5.6 Sol (high) · proprietary
88.8%
11Claude Fable 5 (high) · proprietary
87.8%
12GPT 5.6 Sol Max · proprietary
87.0%
13Claude Opus 5 · proprietary
86.3%
14GPT 5.5 (xhigh) · proprietary
84.9%
15GPT 5.5 (high) · proprietary
83.9%
16Claude Opus 4.8 (xhigh) · proprietary
82.9%
17Kimi K3 · 2779.9B
82.6%
18GPT 5.3 Codex · proprietary
79.3%
19GPT 5.6 Terra (high) · proprietary
78.3%
20Claude Opus 4.6 (high) · proprietary
78.0%
21Claude Opus 4.6 (unspecified) · proprietary
77.9%
22GPT 5.3 Codex (xhigh) · proprietary
77.9%
23GPT 5.4 (Mar 05, 2026, xhigh) · proprietary
77.7%
24Claude Opus 4.7 (high) · proprietary
76.4%
25Claude Opus 4.7 · proprietary
76.4%
26Claude Opus 4.7 (unspecified) · proprietary
76.4%
27Claude Opus 4.8 (medium) · proprietary
76.0%
28Claude Opus 4.7 Max · proprietary
75.4%
29GLM 5.3 · 753.3B
75.4%
30GPT 5.2 (Dec 11, 2025, xhigh) · proprietary
72.2%
31Gemini 3.1 Pro Preview · proprietary
72.1%
32Claude Opus 4.8 None · proprietary
70.5%
33GLM 5.2 · 753.3B
70.1%
34Gemini 3 Pro Preview · proprietary
69.9%
35Claude Sonnet 5 (high) · proprietary
68.8%
36Grok 4.6 (high) · proprietary
67.3%
37GPT 5.5 None · proprietary
67.2%
38DeepSeek V4 Pro 0813 · 1650.5B
66.2%
39Claude Sonnet 4.6 (medium) · proprietary
66.1%
40Claude Opus 4.6 · proprietary
65.9%
41Claude Opus 4.5 (Nov 01, 2025, 16K) · proprietary
63.7%
42GPT 5.2 (Dec 11, 2025, medium) · proprietary
63.4%
43DeepSeek V4 Flash 0731 · 304.2B
63.0%
44Gemini 3.5 Flash (high) · proprietary
62.6%
45Gemini 3 Flash Preview · proprietary
61.6%
46GPT 5.6 Luna (high) · proprietary
60.9%
47GPT 5.1 (Nov 13, 2025, high) · proprietary
60.8%
48GPT 5 (Aug 07, 2025, high) · proprietary
60.7%
49GPT 5 Pro (Oct 06, 2025, high) · proprietary
60.4%
50GPT 5.4 Mini (Mar 17, 2026, high) · proprietary
60.3%
51Muse Spark 1.2 (xhigh) · proprietary
60.3%
52O3 Pro (Jun 10, 2025, high) · proprietary
58.2%
53GPT 5.4 2026 03.05 None · proprietary
57.4%
54GPT 5.4 Pro 2026 03.05 None · proprietary
57.4%
55GLM 5.1 · 753.9B
57.1%
56Gemini 3.6 Flash (high) · proprietary
56.1%
57Kimi K2.6 · 1026.9B
55.9%
58GPT 5 Codex · proprietary
54.5%
59GPT 5 Codex (high) · proprietary
54.5%
60Kimi K2.7 Code · 1026.9B
54.1%
61Gemini 2.5 Pro (16K) · proprietary
54.0%
62GPT 5 Mini (Aug 07, 2025, high) · proprietary
52.7%
63O4 Mini (Apr 16, 2025, high) · proprietary
52.6%
64O3 (Apr 16, 2025, high) · proprietary
52.4%
65Gemma 4 31B IT · 31.3B
52.3%
66Grok 4.20 · proprietary
52.3%
67Gemini 3.1 Flash Lite · proprietary
52.2%
68Grok 4.3 (unspecified) · proprietary
49.9%
69GPT 5.2 (Dec 11, 2025, low) · proprietary
49.6%
70GPT 5.2 2025 12.11 None · proprietary
49.6%
71GPT 5.4 Nano (Mar 17, 2026, high) · proprietary
49.2%
72DeepSeek V4 Pro · 1598.8B
48.9%
73GLM 5 · 753.9B
48.2%
74GPT OSS 120B · 116.8B
48.2%
75Claude Sonnet 4.5 (Sep 29, 2025, 16K) · proprietary
47.7%
76O1 Preview (Sep 12, 2024) · proprietary
47.6%
77DeepSeek V3.2 Speciale · 685.4B
46.7%
78Claude Sonnet 4.5 (Sep 29, 2025) · proprietary
46.7%
79Grok 4.5 (unspecified) · proprietary
46.4%
80Claude Sonnet 4 (May 14, 2025, 16K) · proprietary
46.1%
81O1 (Dec 17, 2024, high) · proprietary
46.1%
82Claude Opus 4.1 (Aug 05, 2025, 16K) · proprietary
45.9%
83Grok 4 (Jul 09) · proprietary
45.7%
84DeepSeek V4 Flash · 290.9B
45.6%
85Kimi K2.5 · 1026.9B
45.6%
86Claude Haiku 4.5 (Oct 01, 2025) · proprietary
45.4%
87Claude Haiku 4.5 (Oct 01, 2025, 16K) · proprietary
44.1%
88Claude Sonnet 4 (May 14, 2025) · proprietary
43.9%
89Claude Opus 4 (May 14, 2025, 16K) · proprietary
43.7%
90Mistral Medium 2604 · proprietary
43.7%
91O3 Mini (Jan 31, 2025, high) · proprietary
43.7%
92NVIDIA Nemotron 3 Ultra 550B A55B BF16 · 560.5B
43.5%
93Mercury 2 · proprietary
43.2%
94Grok 4 Fast · proprietary
42.9%
95Kimi K2 Thinking · 1026.4B
42.8%
96Grok 3 Mini (high) · proprietary
42.6%
97Grok 3 Mini Beta (high) · proprietary
42.6%
98Gemini 2.5 Flash Preview (Sep 2025, 16K) · proprietary
41.9%
99DeepSeek R1 0528 · 684.5B
41.6%
100Qwen3 Coder 480B A35B Instruct · 480.2B
41.2%
101Qwen3 235B A22B Thinking 2507 · 235.1B
41.0%
102Gemini 2.5 Flash Preview (16K thinking) (Apr 17) · proprietary
40.9%
103Gemini 2.5 Flash Preview (May 20, 16K) · proprietary
40.9%
104GPT OSS 20B · 20.9B
40.9%
105GLM 4.5 · 358.3B
40.6%
106Claude 3.5 Sonnet (Oct 22, 2024) · proprietary
40.0%
107GPT 5 Chat · proprietary
39.8%
108Qwen3.5 27B · 27.8B
39.5%
109DeepSeek V3.2 Exp · 685.4B
39.5%
110GPT 4.5 Preview (Feb 27, 2025) · proprietary
39.4%
111Kimi K2 Instruct · 1026.4B
39.4%
112GPT 4.1 (Apr 14, 2025) · proprietary
39.0%
113Gemini 3.5 Flash Lite (high) · proprietary
39.0%
114Qwen3 235B A22B Instruct 2507 · 235.1B
38.7%
115DeepSeek V3.1 · 684.5B
38.4%
116GPT 5 Nano (Aug 07, 2025, high) · proprietary
38.1%
117NVIDIA Nemotron 3 Super 120B A12B BF16 · 123.6B
38.0%
118GPT 5.4 Nano 2026 03.17 None · proprietary
38.0%
119GPT 5.4 Mini 2026 03.17 None · proprietary
37.9%
120GPT 4.1 Mini (Apr 14, 2025) · proprietary
37.6%
121Qwen3 235B A22B · 235.1B
37.3%
122Grok 3 · proprietary
37.2%
123MiniMax M2.7 · 228.7B
37.0%
124Kimi K2 Instruct 0905 · 1026.5B
36.7%
125DeepSeek R1 · 684.5B
36.5%
126O1 Mini (Sep 12, 2024, medium) · proprietary
36.3%
127DeepSeek v3 0324 · 684.5B
36.1%
128Gemini 2.5 Flash Lite Preview (Jun 17, 16K) · proprietary
35.2%
129Gemini 2.5 Flash Lite Preview Thinking (Jun 17) · proprietary
35.2%
130Gemma 4 26B A4B IT · 25.8B
35.2%
131Grok Code Fast 1 · proprietary
35.1%
132Qwen3.6 35B A3B · 36.0B
34.5%
133Qwen3 Coder Next · 79.7B
34.4%
134Mistral Medium 2508 · proprietary
33.1%
135Inkling · 952.4B
32.3%
136Claude 3.5 Sonnet (Jun 20, 2024) · proprietary
31.0%
137Claude 3.5 Haiku (Oct 22, 2024) · proprietary
30.7%
138Qwen3 30B A3B · 30.5B
29.8%
139GPT 5 Nano (Aug 07, 2025, low) · proprietary
25.9%
140Gemini 2.0 Flash 001 · proprietary
25.8%
141GPT 4o (Nov 20, 2024) · proprietary
25.1%
142Gemini 1.5 Flash 002 · proprietary
24.9%
143Llama 4 Maverick 17B 128E Instruct · 401.6B
24.5%
144Grok 2 (Dec 12) · proprietary
22.2%
145Gemini 1.5 Pro 002 · proprietary
22.2%
146Llama 3.1 405B Instruct · 405.9B
21.4%
147Claude 3 Opus (Feb 29, 2024) · proprietary
19.2%
148GPT 4.1 Nano (Apr 14, 2025) · proprietary
19.0%
149GPT 4 Turbo (Apr 09, 2024) · proprietary
18.0%
150Qwen2.5 72B Instruct · 72.7B
16.0%
151Llama 3.3 70B Instruct · 70.6B
14.4%
152GPT 4 (Jun 13) · proprietary
12.4%
153GPT 4o Mini (Jul 18, 2024) · proprietary
11.8%
154Qwen2 72B Instruct · 72.7B
11.3%
155Claude 3 Sonnet (Feb 29, 2024) · proprietary
10.2%
156Claude 3 Haiku (Mar 07, 2024) · proprietary
9.8%
157Llama 3.1 70B Instruct · 70.6B
9.0%
158Claude 2.1 · proprietary
7.1%
159GPT 3.5 Turbo (Jan 25) · proprietary
3.5%
160Mixtral 8x22B Instruct v0.1 · 140.6B
3.2%
161Llama 3.1 8B Instruct · 8.0B
1.7%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

10B100B1Tmodel size (log scale) →82.6%1.7%GLM 5.2 · 753B · 70.1%DeepSeek V4 Pro 0813 · 1.7T · 66.2%GLM 5.1 · 754B · 57.1%Kimi K2.6 · 1T · 55.9%Kimi K2.7 Code · 1T · 54.1%DeepSeek V4 Pro · 1.6T · 48.9%GPT OSS 120B · 117B · 48.2%GLM 5 · 754B · 48.2%DeepSeek V3.2 Speciale · 685B · 46.7%DeepSeek V4 Flash · 291B · 45.6%Kimi K2.5 · 1T · 45.6%NVIDIA Nemotron 3 Ultra 550B A55B BF16 · 561B · 43.5%Kimi K2 Thinking · 1T · 42.8%DeepSeek R1 0528 · 685B · 41.6%Qwen3 Coder 480B A35B Instruct · 480B · 41.2%Qwen3 235B A22B Thinking 2507 · 235B · 41.0%GLM 4.5 · 358B · 40.6%Qwen3.5 27B · 28B · 39.5%DeepSeek V3.2 Exp · 685B · 39.5%Kimi K2 Instruct · 1T · 39.4%Qwen3 235B A22B Instruct 2507 · 235B · 38.7%DeepSeek V3.1 · 685B · 38.4%NVIDIA Nemotron 3 Super 120B A12B BF16 · 124B · 38.0%Qwen3 235B A22B · 235B · 37.3%MiniMax M2.7 · 229B · 37.0%Kimi K2 Instruct 0905 · 1T · 36.7%DeepSeek R1 · 684B · 36.5%DeepSeek v3 0324 · 685B · 36.1%Gemma 4 26B A4B IT · 26B · 35.2%Qwen3.6 35B A3B · 36B · 34.5%Qwen3 Coder Next · 80B · 34.4%Inkling · 952B · 32.3%Qwen3 30B A3B · 31B · 29.8%Llama 4 Maverick 17B 128E Instruct · 402B · 24.5%Llama 3.1 405B Instruct · 406B · 21.4%Qwen2.5 72B Instruct · 73B · 16.0%Llama 3.3 70B Instruct · 71B · 14.4%Qwen2 72B Instruct · 73B · 11.3%Llama 3.1 70B Instruct · 71B · 9.0%Mixtral 8x22B Instruct v0.1 · 141B · 3.2%Llama 3.1 8B Instruct · 8B · 1.7%Llama 3.1 8B InstructGPT OSS 20B · 21B · 40.9%GPT OSS 20BGemma 4 31B IT · 31B · 52.3%Gemma 4 31B ITDeepSeek V4 Flash 0731 · 304B · 63.0%DeepSeek V4 Flash 0731GLM 5.3 · 753B · 75.4%GLM 5.3Kimi K3 · 2.8T · 82.6%Kimi K3
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Llama 3.1 8B Instruct, 8B, score 1.7% — on the efficiency frontier (best score at its size or smaller).
  • GPT OSS 20B, 21B, score 40.9% — on the efficiency frontier (best score at its size or smaller).
  • Gemma 4 31B IT, 31B, score 52.3% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek V4 Flash 0731, 304B, score 63.0% — on the efficiency frontier (best score at its size or smaller).
  • GLM 5.3, 753B, score 75.4% — on the efficiency frontier (best score at its size or smaller).
  • Kimi K3, 2.8T, score 82.6% — on the efficiency frontier (best score at its size or smaller).

WeirdML: frequently asked questions

What is the best open LLM on WeirdML?
Kimi K3 is the top open model on WeirdML, scoring 82.6%. Among all models tested — including proprietary ones — it ranks #17. The top model overall is GPT 6 Astra Promax (OpenAI) at 93.6%.
What's the best WeirdML model you can run on a 24 GB GPU?
Gemma 4 31B IT is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 17 GB), scoring 52.3% on WeirdML.
What's the best WeirdML model you can run on a 12 GB GPU?
GPT OSS 20B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 12 GB), scoring 40.9% on WeirdML.
Can open models match proprietary models on WeirdML?
Not quite on WeirdML: the strongest proprietary model (GPT 6 Astra Promax) scores 93.6%, ahead of the best open model (Kimi K3) at 82.6% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.