Coding

Surface Evolver Bench Leaderboard

Surface Evolver Bench tests whether a model can write a correct simulation file for Surface Evolver, a physics tool that models surfaces shaped by tension, gravity and volume constraints, defining a 3D structure in its domain-specific format from only a few examples. Scores compare the resulting geometry — volume, area and energy — against reference solutions; the benchmark was built by an independent developer (GitHub: yhenon).

Source: epoch13 open models ranked+14 proprietaryData through Sep 2026

All models ranked on Surface Evolver Bench

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1Claude Fable 5 (high) · proprietary
95.0%
2Kimi K3 · 2779.9B
95.0%
3GPT 5.6 Sol (xhigh) · proprietary
93.1%
4GPT 5.5 (high) · proprietary
88.1%
5Claude Opus 4.8 (high) · proprietary
87.5%
6GPT 5.6 Terra (xhigh) · proprietary
83.8%
7GPT 5.5 (medium) · proprietary
81.3%
8Gemini 3.8 Flash (high) · proprietary
76.9%
9Grok 4.5 (high) · proprietary
74.4%
10Claude Opus 4.8 None · proprietary
68.1%
11GPT 5.6 Luna (medium) · proprietary
61.9%
12Claude Sonnet 5 (medium) · proprietary
60.0%
13Gemini 3.5 Flash (medium) · proprietary
58.1%
14GLM 5.2 · 753.3B
55.6%
15MiniMax M3 · 427.0B
55.0%
16GLM 5.3 Flash · 321.3B
52.5%
17Muse Spark 1.1 (high) · proprietary
52.5%
18Kimi K2.7 Code · 1026.9B
48.8%
19DeepSeek V4.1 Flash · 763.2B
46.3%
20Qwen3.8 27B · 27.8B
45.0%
21Qwen3.6 35B A3B · 36.0B
44.4%
22DeepSeek V4 Pro · 1598.8B
40.0%
23Gemma 4 31B IT · 31.3B
30.6%
24Mistral Medium 2604 · proprietary
26.9%
25GPT OSS 120B · 116.8B
25.0%
26Laguna M.1 · 225.8B
15.6%
27Trinity Large Thinking · 398.6B
15.6%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

100B1Tmodel size (log scale) →95.0%15.6%Kimi K2.7 Code · 1T · 48.8%DeepSeek V4.1 Flash · 763B · 46.3%Qwen3.6 35B A3B · 36B · 44.4%DeepSeek V4 Pro · 1.6T · 40.0%Gemma 4 31B IT · 31B · 30.6%GPT OSS 120B · 117B · 25.0%Trinity Large Thinking · 399B · 15.6%Laguna M.1 · 226B · 15.6%Qwen3.8 27B · 28B · 45.0%Qwen3.8 27BGLM 5.3 Flash · 321B · 52.5%MiniMax M3 · 427B · 55.0%MiniMax M3GLM 5.2 · 753B · 55.6%GLM 5.2Kimi K3 · 2.8T · 95.0%Kimi K3
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Qwen3.8 27B, 28B, score 45.0% — on the efficiency frontier (best score at its size or smaller).
  • GLM 5.3 Flash, 321B, score 52.5% — on the efficiency frontier (best score at its size or smaller).
  • MiniMax M3, 427B, score 55.0% — on the efficiency frontier (best score at its size or smaller).
  • GLM 5.2, 753B, score 55.6% — on the efficiency frontier (best score at its size or smaller).
  • Kimi K3, 2.8T, score 95.0% — on the efficiency frontier (best score at its size or smaller).

Surface Evolver Bench: frequently asked questions

What is the best open LLM on Surface Evolver Bench?
Kimi K3 is the top open model on Surface Evolver Bench, scoring 95.0%. Among all models tested — including proprietary ones — it ranks #1. That puts it ahead of every proprietary model we track, including Claude Fable 5 (high) (Anthropic) at 95.0%.
What's the best Surface Evolver Bench model you can run on a 24 GB GPU?
Qwen3.8 27B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 15 GB), scoring 45.0% on Surface Evolver Bench.
Can open models match proprietary models on Surface Evolver Bench?
Yes — the best open model (Kimi K3, 95.0%) matches or beats every proprietary model we track on Surface Evolver Bench.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.

Surface Evolver Bench Leaderboard — LLM Scores | llmrun