BenchLeader
Zhipu AIOpen weights: zai-org/GLM-5.1Reasoning modelAuto-detected

GLM-5.1

Best configuration ranks #92 of 610 on the BenchLeader Index at 60.1 ±2.9. Last measured 8 Sept 2026. Released 7 Apr 2026.

Blended price
$2.15/M
$1.40 in · $4.40 out
Output speed
60 tok/s
First answer
66 s
first token 1.35 s
Context
200k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index60
  2. Reasoning61
  3. Coding57
  4. Agents & tools57
  5. Maths53
  6. Knowledge57
  7. Instruction following74
  8. Human preference66
  9. Long context63
  10. Composite62

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning52.6#24438 tok/s1.97 s$0.0019Agents & tools 65 · Composite 60 · Instruction following 53 · Knowledge 53 · Long context 53 · Maths 41
defaultbest60.1#9260 tok/s66 s$0.0019Agents & tools 57 · Coding 57 · Composite 62 · Human preference 66 · Instruction following 74 · Knowledge 57 · Long context 63 · Maths 53 · Reasoning 61

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoningdefaultSourceTrend
GPQA Diamond89.9%#39Epoch AI Benchmarking Hub
SimpleBench55.1%#39SimpleBench
LMArena Hard Prompts1489#42LMArena
GPQA Diamond (AA)not in index83.9%#15086.8%#106Artificial Analysis
Humanity's Last Exam (AA)not in index27.9%#12130.1%#103Artificial Analysis
GPQA Diamond (Vals)not in index84.5%#50Vals AI

Coding

Benchmarkno reasoningdefaultSourceTrend
SWE-bench Verified (Epoch)74.2%#16Epoch AI Benchmarking Hub
SciCode43.8%#90SciCode
WeirdML57.1%#52WeirdML
LMArena Coding1514#37LMArena
LMArena WebDev1508#39LMArena
SciCode (AA)not in index44.8%#94Artificial Analysis
LiveCodeBench81.4%#60Vals AI
SWE-bench (Vals)not in index76.4%#35Vals AI

Agents & tools

Benchmarkno reasoningdefaultSourceTrend
Terminal-Bench Hard35.6%#7243.2%#39Artificial Analysis
τ²-Bench Telecom (AA)not in index97.1%#1397.7%#9Artificial Analysis
Terminal-Bench 2.1 (Vals)56.9%#38Vals AI
MCP Atlas75.6%#15Scale AI SEAL

Knowledge

Benchmarkno reasoningdefaultSourceTrend
SimpleQA Verified34.0%#46Epoch AI Benchmarking Hub
AA-Omniscience-22.4#1890.8#88Artificial Analysis
MMLU-Pro86.9%#36Vals AI
LegalBench84.4%#30Vals AI
CorpFin64.5%#37Vals AI
TaxEval71.2%#79Vals AI

Instruction following

Benchmarkno reasoningdefaultSourceTrend
IFBench52.0%#16176.3%#22Artificial Analysis

Human preference

Benchmarkno reasoningdefaultSourceTrend
LMArena Text1466#44LMArena

Long context

Benchmarkno reasoningdefaultSourceTrend
AA-LCR53.3%#24173.7%#125Artificial Analysis

Composite

Benchmarkno reasoningdefaultSourceTrend
Epoch Capabilities Indexnot in index149.7#46Epoch AI Benchmarking Hub
AA Intelligence Index24.2#12526.4#106Artificial Analysis

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Friendli99 tok/s0.30 s$1.40$4.40203k
Crusoe78 tok/s0.56 s$1.20$4.40203kfp8
Baidu Qianfan60 tok/s0.97 s$0.965$3.03203kfp8
AtlasCloud58 tok/s1.21 s$1.26$3.96203kfp8
NovitaAI57 tok/s2.06 s$1.38$4.40205kfp8
GMICloud55 tok/s1.51 s$1.40$4.40203kfp8
DeepInfra53 tok/s0.91 s$1.05$3.50203kfp4
Venice51 tok/s1.30 s$1.54$4.84200kfp8
Alibaba Cloud Int.50 tok/s1.52 s$1.33$4.18203kfp8
StreamLake37 tok/s1.35 s$0.966$3.04200kfp8
SiliconFlow28 tok/s1.66 s$1.19$3.74205kfp8
Nebius Token Factory28 tok/s0.96 s$1.40$4.40203kfp8
Chutes24 tok/s5.73 s$0.980$3.08203kfp8
Phala21 tok/s3.93 s$1.21$4.20203k
Z.ai20 tok/s7.91 s$1.40$4.40203kfp8

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$1.21$2.42$3.63$4.84May 26Jun 26Jul 26Aug 26
input outputnow $0.965 in · $3.03 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.260 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0019$0.00151.2 min
Summarise a 30-page report12,000 / 600$0.019$0.00921.3 min
Code edit6,000 / 1,500$0.015$0.00991.5 min
Agentic coding session60,000 / 4,000$0.102$0.0502.2 min
Structured extraction2,000 / 200$0.0037$0.00201.1 min

See also

Data as of 9 Sept 2026. Compare these configurations.