Coding

Surface Evolver Bench Leaderboard

Surface Evolver Bench tests whether a model can write a correct simulation file for Surface Evolver, a physics tool that models surfaces shaped by tension, gravity and volume constraints, defining a 3D structure in its domain-specific format from only a few examples. Scores compare the resulting geometry — volume, area and energy — against reference solutions; the benchmark was built by an independent developer (GitHub: yhenon).

Source: epoch13 open models ranked+14 proprietaryData through Sep 2026

Open models ranked on Surface Evolver Bench

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 2Kimi K3 · 2779.9B
95.0%
2 / 14GLM 5.2 · 753.3B
55.6%
3 / 15MiniMax M3 · 427.0B
55.0%
4 / 16GLM 5.3 Flash · 321.3B
52.5%
5 / 18Kimi K2.7 Code · 1026.9B
48.8%
6 / 19DeepSeek V4.1 Flash · 763.2B
46.3%
7 / 20Qwen3.8 27B · 27.8B
45.0%
8 / 21Qwen3.6 35B A3B · 36.0B
44.4%
9 / 22DeepSeek V4 Pro · 1598.8B
40.0%
10 / 23Gemma 4 31B IT · 31.3B
30.6%
11 / 25GPT OSS 120B · 116.8B
25.0%
12 / 26Laguna M.1 · 225.8B
15.6%
13 / 27Trinity Large Thinking · 398.6B
15.6%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

100B1Tmodel size (log scale) →95.0%15.6%Kimi K2.7 Code · 1T · 48.8%DeepSeek V4.1 Flash · 763B · 46.3%Qwen3.6 35B A3B · 36B · 44.4%DeepSeek V4 Pro · 1.6T · 40.0%Gemma 4 31B IT · 31B · 30.6%GPT OSS 120B · 117B · 25.0%Trinity Large Thinking · 399B · 15.6%Laguna M.1 · 226B · 15.6%Qwen3.8 27B · 28B · 45.0%Qwen3.8 27BGLM 5.3 Flash · 321B · 52.5%MiniMax M3 · 427B · 55.0%MiniMax M3GLM 5.2 · 753B · 55.6%GLM 5.2Kimi K3 · 2.8T · 95.0%Kimi K3
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Qwen3.8 27B, 28B, score 45.0% — on the efficiency frontier (best score at its size or smaller).
  • GLM 5.3 Flash, 321B, score 52.5% — on the efficiency frontier (best score at its size or smaller).
  • MiniMax M3, 427B, score 55.0% — on the efficiency frontier (best score at its size or smaller).
  • GLM 5.2, 753B, score 55.6% — on the efficiency frontier (best score at its size or smaller).
  • Kimi K3, 2.8T, score 95.0% — on the efficiency frontier (best score at its size or smaller).

Surface Evolver Bench: frequently asked questions

What is the best open LLM on Surface Evolver Bench?
Kimi K3 is the top open model on Surface Evolver Bench, scoring 95.0%. Among all models tested — including proprietary ones — it ranks #1. That puts it ahead of every proprietary model we track, including Claude Fable 5 (high) (Anthropic) at 95.0%.
What's the best Surface Evolver Bench model you can run on a 24 GB GPU?
Qwen3.8 27B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 15 GB), scoring 45.0% on Surface Evolver Bench.
Can open models match proprietary models on Surface Evolver Bench?
Yes — the best open model (Kimi K3, 95.0%) matches or beats every proprietary model we track on Surface Evolver Bench.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.