Surface Evolver Bench Leaderboard
Surface Evolver Bench tests whether a model can write a correct simulation file for Surface Evolver, a physics tool that models surfaces shaped by tension, gravity and volume constraints, defining a 3D structure in its domain-specific format from only a few examples. Scores compare the resulting geometry — volume, area and energy — against reference solutions; the benchmark was built by an independent developer (GitHub: yhenon).
Source: epoch13 open models ranked+14 proprietaryData through Sep 2026
All models ranked on Surface Evolver Bench
Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.
| # | Model | Score |
|---|---|---|
| 1 | Claude Fable 5 (high) · proprietary | 95.0% |
| 2 | Kimi K3 · 2779.9B | 95.0% |
| 3 | GPT 5.6 Sol (xhigh) · proprietary | 93.1% |
| 4 | GPT 5.5 (high) · proprietary | 88.1% |
| 5 | Claude Opus 4.8 (high) · proprietary | 87.5% |
| 6 | GPT 5.6 Terra (xhigh) · proprietary | 83.8% |
| 7 | GPT 5.5 (medium) · proprietary | 81.3% |
| 8 | Gemini 3.8 Flash (high) · proprietary | 76.9% |
| 9 | Grok 4.5 (high) · proprietary | 74.4% |
| 10 | Claude Opus 4.8 None · proprietary | 68.1% |
| 11 | GPT 5.6 Luna (medium) · proprietary | 61.9% |
| 12 | Claude Sonnet 5 (medium) · proprietary | 60.0% |
| 13 | Gemini 3.5 Flash (medium) · proprietary | 58.1% |
| 14 | GLM 5.2 · 753.3B | 55.6% |
| 15 | MiniMax M3 · 427.0B | 55.0% |
| 16 | GLM 5.3 Flash · 321.3B | 52.5% |
| 17 | Muse Spark 1.1 (high) · proprietary | 52.5% |
| 18 | Kimi K2.7 Code · 1026.9B | 48.8% |
| 19 | DeepSeek V4.1 Flash · 763.2B | 46.3% |
| 20 | Qwen3.8 27B · 27.8B | 45.0% |
| 21 | Qwen3.6 35B A3B · 36.0B | 44.4% |
| 22 | DeepSeek V4 Pro · 1598.8B | 40.0% |
| 23 | Gemma 4 31B IT · 31.3B | 30.6% |
| 24 | Mistral Medium 2604 · proprietary | 26.9% |
| 25 | GPT OSS 120B · 116.8B | 25.0% |
| 26 | Laguna M.1 · 225.8B | 15.6% |
| 27 | Trinity Large Thinking · 398.6B | 15.6% |
Score vs model size
Which models give the most quality for their size — the ones worth running locally.
- Qwen3.8 27B, 28B, score 45.0% — on the efficiency frontier (best score at its size or smaller).
- GLM 5.3 Flash, 321B, score 52.5% — on the efficiency frontier (best score at its size or smaller).
- MiniMax M3, 427B, score 55.0% — on the efficiency frontier (best score at its size or smaller).
- GLM 5.2, 753B, score 55.6% — on the efficiency frontier (best score at its size or smaller).
- Kimi K3, 2.8T, score 95.0% — on the efficiency frontier (best score at its size or smaller).
Surface Evolver Bench: frequently asked questions
- What is the best open LLM on Surface Evolver Bench?
- Kimi K3 is the top open model on Surface Evolver Bench, scoring 95.0%. Among all models tested — including proprietary ones — it ranks #1. That puts it ahead of every proprietary model we track, including Claude Fable 5 (high) (Anthropic) at 95.0%.
- What's the best Surface Evolver Bench model you can run on a 24 GB GPU?
- Qwen3.8 27B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 15 GB), scoring 45.0% on Surface Evolver Bench.
- Can open models match proprietary models on Surface Evolver Bench?
- Yes — the best open model (Kimi K3, 95.0%) matches or beats every proprietary model we track on Surface Evolver Bench.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.