BenchLeader
Zhipu AIOpen weights: zai-org/GLM-5.2Reasoning modelAuto-detected

GLM-5.2

Best configuration ranks #138 of 610 on the BenchLeader Index at 57.6 ±3.7 (max reasoning effort). Last measured 8 Sept 2026. Released 16 Jun 2026.

Blended price
$1.79/M
$1.10 in · $3.85 out
Output speed
63 tok/s
First answer
35 s
first token 1.22 s
Context
1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index58
  2. Reasoning66
  3. Coding64
  4. Agents & tools57
  5. Maths56
  6. Knowledge45
  7. Human preference67

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning50.3#287105 tok/s2.08 s$0.0019Coding 39 · Composite 57 · Knowledge 61 · Long context 47 · Maths 46 · Reasoning 52
low63 tok/s35 s$0.0016Maths 58 · Reasoning 64
high63 tok/s35 s$0.0016Coding 62
maxbest57.6#13863 tok/s35 s$0.0016Agents & tools 57 · Coding 64 · Human preference 67 · Knowledge 45 · Maths 56 · Reasoning 66
thinking63 tok/s35 s$0.0016Reasoning 55
default57.4#13963 tok/s35 s$0.0019Agents & tools 63 · Coding 49 · Composite 62 · Human preference 54 · Instruction following 71 · Knowledge 61 · Long context 66 · Maths 47 · Reasoning 53

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoninglowhighmaxthinkingdefaultSourceTrend
GPQA Diamond71.2%#15487.9%#5391.9%#23Epoch AI Benchmarking Hub
SimpleBench58.8%#32SimpleBench
LMArena Hard Prompts1492#38LMArena
LiveBench Reasoningnot in index78.6%#41LiveBench
GPQA Diamond (AA)not in index68.6%#29889.5%#74Artificial Analysis
Humanity's Last Exam (AA)not in index9.8%#26241.1%#49Artificial Analysis
GPQA Diamond (Vals)not in index85.6%#44Vals AI
Kagi LLM Benchmark60.0%#5062.6%#44Kagi LLM Benchmark
ARC-AGI-177.0%#87ARC Prize
ARC-AGI-222.8%#93ARC Prize

Coding

Benchmarkno reasoninglowhighmaxthinkingdefaultSourceTrend
SWE-bench Verified (Epoch)78.7%#5Epoch AI Benchmarking Hub
SciCode36.1%#12550.5%#56SciCode
WeirdML67.3%#3370.1%#30WeirdML
FrontierCode24.5%#22Cognition
LMArena Coding1509#47LMArena
LMArena WebDev1589#16LMArena
LiveBench Codingnot in index79.7%#14LiveBench
SciCode (AA)not in index51.2%#61Artificial Analysis
LiveCodeBench69.5%#89Vals AI
SWE-bench (Vals)not in index82.8%#18Vals AI

Agents & tools

Benchmarkno reasoninglowhighmaxthinkingdefaultSourceTrend
APEX-Agents35.6%#17Mercor
LMArena Agent4.6#11LMArena
LiveBench Agentic Codingnot in index51.8%#28LiveBench
Terminal-Bench Hard50.8%#22Artificial Analysis
τ²-Bench Telecom (AA)not in index99.1%#1Artificial Analysis
Terminal-Bench 2.1 (Vals)67.8%#25Vals AI
MCP Atlas77.8%#13Scale AI SEAL
HiL-Bench43.7%#4Scale AI SEAL

Maths

Benchmarkno reasoninglowhighmaxthinkingdefaultSourceTrend
FrontierMath Tiers 1–342.5%#4554.7%#3659.2%#28Epoch AI Benchmarking Hub
FrontierMath Tier 429.3%#28Epoch AI Benchmarking Hub
OTIS Mock AIME28.9%#19175.6%#11186.4%#78Epoch AI Benchmarking Hub
ProofBench35.0%#27Vals AI
LiveBench Mathematicsnot in index89.8%#25LiveBench
AIME 202690.0%#27MathArena
HMMT February 202692.4%#13MathArena
MathArena Apex19.8%#15MathArena

Knowledge

Benchmarkno reasoninglowhighmaxthinkingdefaultSourceTrend
SimpleQA Verified34.2%#45Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index73.7%#31LiveBench
AA-Omniscience-6.6#1194.4#75Artificial Analysis
MMLU-Pro86.7%#38Vals AI
LegalBench84.1%#36Vals AI
CorpFin66.1%#21Vals AI
TaxEval73.3%#49Vals AI

Instruction following

Benchmarkno reasoninglowhighmaxthinkingdefaultSourceTrend
LiveBench Languagenot in index76.2%#35LiveBench
IFBench73.3%#41Artificial Analysis

Human preference

Benchmarkno reasoninglowhighmaxthinkingdefaultSourceTrend
LMArena Text1472#36LMArena
EQ-Bench 41222#13EQ-Bench

Long context

Benchmarkno reasoninglowhighmaxthinkingdefaultSourceTrend
AA-LCR42.3%#28078.3%#77Artificial Analysis

Composite

Benchmarkno reasoninglowhighmaxthinkingdefaultSourceTrend
Epoch Capabilities Indexnot in index151.9#37Epoch AI Benchmarking Hub
LiveBench73.2%#36LiveBench
AA Intelligence Index22.4#14538.6#43Artificial Analysis
Vals Indexnot in index53.1#23Vals AI

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Decart Fast225 tok/s1.06 s$2.10$6.601.0Mfp4
Fireworks Fast (US)136 tok/s0.91 s$2.10$6.601.0M
Baseten Fast118 tok/s1.95 s$2.10$6.601.0Mfp8
Fireworks Fast106 tok/s0.72 s$2.10$6.601.0M
Crusoe102 tok/s0.75 s$1.40$4.401.0Mfp8
Mistral99 tok/s0.73 s$1.40$4.401.0M
Friendli99 tok/s0.43 s$1.40$4.401.0M
Mistral (ZDR)97 tok/s1.01 s$1.40$4.401.0M
Mistral (EU)80 tok/s1.22 s$1.54$4.841.0M
Alibaba Fast80 tok/s0.92 s$2.31$7.261.0Mfp8
Baseten74 tok/s1.58 s$1.40$4.401.0Mfp8
AtlasCloud67 tok/s1.07 s$0.938$2.951.0Mfp8
Baidu Qianfan (fp4)65 tok/s0.92 s$0.682$2.391.0Mfp4
Baseten (US)61 tok/s2.23 s$1.40$4.401.0Mfp8
Baseten Fast (US)61 tok/s2.47 s$2.10$6.601.0Mfp8
Baidu Qianfan (fp8)60 tok/s0.91 s$0.487$1.531.0Mfp8
Alibaba Cloud Int.60 tok/s0.96 s$0.966$3.041.0Mfp8
CoreWeave58 tok/s1.46 s$0.760$2.421.0Mfp4
Venice56 tok/s1.89 s$1.40$4.401Mfp8
Phala51 tok/s1.54 s$1.26$3.001.0Mfp8
DigitalOcean48 tok/s0.91 s$0.700$2.20262k
SiliconFlow48 tok/s1.56 s$1.19$3.741.0Mfp8
Z.ai45 tok/s3.33 s$1.40$4.401.0Mfp8
Fireworks44 tok/s2.16 s$1.40$4.401.0M
Together39 tok/s0.63 s$1.40$4.40512k
Parasail39 tok/s1.70 s$1.40$4.40262kfp4
StreamLake36 tok/s3.83 s$0.280$0.8801.0Mfp8
DeepInfra36 tok/s1.03 s$0.487$1.561.0Mfp4
NovitaAI35 tok/s2.27 s$0.683$2.151.0Mfp8
Inceptron31 tok/s1.02 s$1.25$2.991.0Mfp4
Cloudflare23 tok/s4.24 s$1.40$4.40262k
GMICloud19 tok/s2.95 s$1.40$4.401.0Mfp8
Ambient13 tok/s2.91 s$0.600$2.00203kfp8

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$1.21$2.42$3.63$4.84Jun 26Jul 26Aug 26
input outputnow $0.280 in · $0.880 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.275 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0016$0.001340.2 s
Summarise a 30-page report12,000 / 600$0.015$0.008145.0 s
Code edit6,000 / 1,500$0.012$0.008759.3 s
Agentic coding session60,000 / 4,000$0.081$0.0441.7 min
Structured extraction2,000 / 200$0.0030$0.001738.6 s

See also

Data as of 9 Sept 2026. Compare these configurations.