Surface Evolver Bench Leaderboard
Surface Evolver Bench tests whether a model can write a correct simulation file for Surface Evolver, a physics tool that models surfaces shaped by tension, gravity and volume constraints, defining a 3D structure in its domain-specific format from only a few examples. Scores compare the resulting geometry — volume, area and energy — against reference solutions; the benchmark was built by an independent developer (GitHub: yhenon).
Source: epoch13 open models ranked+14 proprietaryData through Sep 2026
Open models ranked on Surface Evolver Bench
# shows rank among open models / rank overall (including proprietary).
| # | Model | Score |
|---|---|---|
| 1 / 2 | Kimi K3 · 2779.9B | 95.0% |
| 2 / 14 | GLM 5.2 · 753.3B | 55.6% |
| 3 / 15 | MiniMax M3 · 427.0B | 55.0% |
| 4 / 16 | GLM 5.3 Flash · 321.3B | 52.5% |
| 5 / 18 | Kimi K2.7 Code · 1026.9B | 48.8% |
| 6 / 19 | DeepSeek V4.1 Flash · 763.2B | 46.3% |
| 7 / 20 | Qwen3.8 27B · 27.8B | 45.0% |
| 8 / 21 | Qwen3.6 35B A3B · 36.0B | 44.4% |
| 9 / 22 | DeepSeek V4 Pro · 1598.8B | 40.0% |
| 10 / 23 | Gemma 4 31B IT · 31.3B | 30.6% |
| 11 / 25 | GPT OSS 120B · 116.8B | 25.0% |
| 12 / 26 | Laguna M.1 · 225.8B | 15.6% |
| 13 / 27 | Trinity Large Thinking · 398.6B | 15.6% |
Score vs model size
Which models give the most quality for their size — the ones worth running locally.
- Qwen3.8 27B, 28B, score 45.0% — on the efficiency frontier (best score at its size or smaller).
- GLM 5.3 Flash, 321B, score 52.5% — on the efficiency frontier (best score at its size or smaller).
- MiniMax M3, 427B, score 55.0% — on the efficiency frontier (best score at its size or smaller).
- GLM 5.2, 753B, score 55.6% — on the efficiency frontier (best score at its size or smaller).
- Kimi K3, 2.8T, score 95.0% — on the efficiency frontier (best score at its size or smaller).
Surface Evolver Bench: frequently asked questions
- What is the best open LLM on Surface Evolver Bench?
- Kimi K3 is the top open model on Surface Evolver Bench, scoring 95.0%. Among all models tested — including proprietary ones — it ranks #1. That puts it ahead of every proprietary model we track, including Claude Fable 5 (high) (Anthropic) at 95.0%.
- What's the best Surface Evolver Bench model you can run on a 24 GB GPU?
- Qwen3.8 27B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 15 GB), scoring 45.0% on Surface Evolver Bench.
- Can open models match proprietary models on Surface Evolver Bench?
- Yes — the best open model (Kimi K3, 95.0%) matches or beats every proprietary model we track on Surface Evolver Bench.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.