BenchLeader

GLM-4.5

Best configuration ranks #289 of 610 on the BenchLeader Index at 50.2 ±3.2. Last measured 2 Sept 2026. Released 28 Jul 2025.

Blended price
$1.00/M
$0.600 in · $2.20 out
Output speed
17 tok/s
First answer
9.78 s
Context
131k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index50
  2. Reasoning48
  3. Coding48
  4. Agents & tools53
  5. Maths51
  6. Knowledge50
  7. Instruction following46
  8. Human preference60
  9. Long context52
  10. Composite45

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
thinking17 tok/s9.78 s$0.0009Coding 44
defaultbest50.2#28917 tok/s9.78 s$0.0009Agents & tools 53 · Coding 48 · Composite 45 · Human preference 60 · Instruction following 46 · Knowledge 50 · Long context 52 · Maths 51 · Reasoning 48

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Coding

BenchmarkthinkingdefaultSource
WeirdML40.6%#100WeirdML
LMArena Coding1455#122LMArena
LiveCodeBench67.5%#93Vals AI
IOI2.9%#52Vals AI
SWE-bench Verified (bash only)54.2%#30SWE-bench
SWE-bench Verified (any scaffold)not in index54.2%#38SWE-bench

Agents & tools

BenchmarkthinkingdefaultSource
Terminal-Bench Hard22.0%#149Artificial Analysis
τ²-Bench Telecom (AA)not in index43.0%#211Artificial Analysis

Maths

BenchmarkthinkingdefaultSource
AIME (Vals)86.7%#31Vals AI
MGSM90.8%#35Vals AI
MathArena Apex1.0%#35MathArena

Knowledge

BenchmarkthinkingdefaultSource
AA-Omniscience-27.4#208Artificial Analysis
MMLU-Pro81.2%#80Vals AI
LegalBench75.6%#104Vals AI
CorpFin61.0%#61Vals AI
TaxEval72.4%#62Vals AI
MedQA90.0%#49Vals AI
MultiNRC17.4%#37Scale AI SEAL

Instruction following

BenchmarkthinkingdefaultSource
IFBench44.1%#213Artificial Analysis

Human preference

BenchmarkthinkingdefaultSource
LMArena Text1411#120LMArena

Long context

BenchmarkthinkingdefaultSource
AA-LCR52.7%#247Artificial Analysis

Composite

BenchmarkthinkingdefaultSource
AA Intelligence Index12.8#257Artificial Analysis

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Z.ai17 tok/s9.78 s$0.600$2.20131kfp8

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.110 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0009$0.000827.4 s
Summarise a 30-page report12,000 / 600$0.0085$0.004145.1 s
Code edit6,000 / 1,500$0.0069$0.00471.6 min
Agentic coding session60,000 / 4,000$0.045$0.0234.1 min
Structured extraction2,000 / 200$0.0016$0.000921.5 s

See also

Data as of 9 Sept 2026. Compare these configurations.