BenchLeader

Surface Evolver Bench

Scientific computing with the Surface Evolver package.

As of 19 Sept 2026, Kimi K3 leads Surface Evolver Bench on BenchLeader with 95.0%, ahead of Claude Fable 5 at 95.0%, across 26 model configurations with a published result.

Published by
Epoch AI Benchmarking Hub
Category
Coding
Index weight
Reference only
Models
26
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

Tasks that require writing correct inputs for the Surface Evolver, a niche scientific tool.

How it is scored

Mean score, as published.

What to keep in mind

Niche tool; small set.

26 of 26
#
1Kimi K3Moonshot AIopen ↗95.0%64.22026-07-16
2Claude Fable 5highAnthropic95.0%64.02026-06-09
3GPT-5.6 SolxhighOpenAI93.1%68.22026-07-09
4Kimi K3maxMoonshot AIopen ↗93.0%67.22026-07-16
5GPT-5.5highOpenAI88.1%67.02026-04-23
6Claude Opus 4.8highAnthropic87.5%62.22026-05-28
7GPT-5.6 TerraxhighOpenAI83.8%64.32026-07-09
8GPT-5.5mediumOpenAI81.3%64.62026-04-23
9Grok 4.5highxAI74.4%61.02026-07-08
10Claude Opus 4.8no reasoningAnthropic68.1%2026-05-28
11GPT-5.6 LunamediumOpenAI61.9%54.02026-07-09
12Claude Sonnet 5mediumAnthropic60.0%2026-06-30
13Gemini 3.5 FlashmediumGoogle58.1%64.22026-05-19
14GLM-5.2highZhipu AIopen ↗55.6%2026-06-16
15MiniMax-M3MiniMaxopen ↗55.0%57.52026-06-01
16GLM-5.3-FlashmaxZhipu AIopen ↗52.5%55.32026-08-20
17Muse Spark 1.1highMeta52.5%2026-07-09
18Kimi K2.7 CodeMoonshot AIopen ↗48.8%55.92026-06-12
19Qwen3.8 27BxhighAlibabaopen ↗45.0%59.22026-08-14
20Qwen3.6 35B-A3BAlibaba44.4%46.72026-04-14
21Deepseek v4 ProhighDeepSeek40.0%60.82026-04-24
22Gemma 4 31BGoogleopen ↗30.6%52.72026-04-02
23Mistral Medium 3.5Mistral AIopen26.9%2026-04-28
24gpt-oss-120bOpenAIopen25.0%46.92025-08-05
25Trinity LargethinkingArcee AIopen15.6%47.62026-04-01
26Laguna M.1Poolsideopen15.6%37.72026-04-28

Cite as: BenchLeader, “Surface Evolver Bench leaderboard”, https://www.benchleader.com/benchmarks/surface_evolver_bench, data as of 19 Sept 2026.

Surface Evolver Bench: questions

What does Surface Evolver Bench measure?
Tasks that require writing correct inputs for the Surface Evolver, a niche scientific tool. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads Surface Evolver Bench?
Kimi K3 leads Surface Evolver Bench with 95.0% as of 19 Sept 2026, ahead of Claude Fable 5 at 95.0%.
How many models have Surface Evolver Bench results?
26 model configurations have a Surface Evolver Bench result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
Who runs Surface Evolver Bench and how often is it updated?
Surface Evolver Bench is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
Does Surface Evolver Bench count toward the BenchLeader Index?
No. Surface Evolver Bench is shown for reference but left out of the composite index.