BenchLeader

GLM-5.3-Flash

Best configuration ranks #74 of 610 on the BenchLeader Index at 61.0 ±6.2. Last measured 8 Sept 2026. Released 20 Aug 2026.

Blended price
$0.119/M
$0.075 in · $0.250 out
Output speed
72 tok/s
First answer
30 s
first token 1.70 s
Context
1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index61
  2. Reasoning67
  3. Coding64
  4. Agents & tools52
  5. Knowledge68
  6. Human preference67
  7. Multimodal65
  8. Long context67
  9. Composite61

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
max55.6#17372 tok/s30 s$0.0001Agents & tools 52 · Coding 58 · Knowledge 59 · Maths 53 · Reasoning 65
defaultbest61.0#7472 tok/s30 s$0.0001Agents & tools 52 · Coding 64 · Composite 61 · Human preference 67 · Knowledge 68 · Long context 67 · Multimodal 65 · Reasoning 67

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkmaxdefaultSourceTrend
GPQA Diamond90.2%#37Epoch AI Benchmarking Hub
LMArena Hard Prompts1496#31LMArena
LiveBench Reasoningnot in index77.6%#43LiveBench
GPQA Diamond (AA)not in index91.2%#49Artificial Analysis
Humanity's Last Exam (AA)not in index39.9%#56Artificial Analysis
GPQA Diamond (Vals)not in index86.4%#41Vals AI

Coding

BenchmarkmaxdefaultSourceTrend
SciCode46.1%#81SciCode
LMArena Coding1534#9LMArena
LMArena WebDev1605#14LMArena
LiveBench Codingnot in index79.0%#17LiveBench
SciCode (AA)not in index51.6%#56Artificial Analysis
LiveCodeBench80.5%#66Vals AI
SWE-bench (Vals)not in index92.0%#10Vals AI

Agents & tools

BenchmarkmaxdefaultSourceTrend
LMArena Agent2#20LMArena
LiveBench Agentic Codingnot in index56.8%#19LiveBench
Terminal-Bench 2.1 (Vals)62.9%#30Vals AI

Knowledge

BenchmarkmaxdefaultSourceTrend
LiveBench Data Analysisnot in index76.4%#24LiveBench
AA-Omniscience7.5#66Artificial Analysis
MMLU-Pro86.1%#47Vals AI
LegalBench83.9%#41Vals AI
TaxEval75.6%#16Vals AI

Instruction following

BenchmarkmaxdefaultSourceTrend
LiveBench Languagenot in index77.3%#32LiveBench

Human preference

BenchmarkmaxdefaultSourceTrend
LMArena Text1474#29LMArena

Multimodal

BenchmarkmaxdefaultSourceTrend
LMArena Vision1296#15LMArena

Long context

BenchmarkmaxdefaultSourceTrend
AA-LCR80.0%#51Artificial Analysis

Composite

BenchmarkmaxdefaultSourceTrend
Epoch Capabilities Indexnot in index151.4#39Epoch AI Benchmarking Hub
LiveBench71.6%#40LiveBench
AA Intelligence Index41.9#30Artificial Analysis
Vals Indexnot in index47.2#30Vals AI

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Morph135 tok/s0.99 s$0.388$1.361.0Mfp8
Crusoe115 tok/s0.95 s$0.150$0.5001.0Mfp4
Baseten102 tok/s0.80 s$0.150$0.5001.0Mfp8
Friendli82 tok/s1.03 s$0.150$0.5001.0M
Modal71 tok/s0.75 s$0.150$0.5001.0Mfp8
Relace70 tok/s0.95 s$0.120$0.4001.0Mfp4
CoreWeave65 tok/s1.11 s$0.150$0.5001.0Mfp8
Together62 tok/s1.09 s$0.150$0.5001.0M
Parasail57 tok/s1.22 s$0.150$0.5001.0Mfp8
Makora55 tok/s1.04 s$0.140$0.4701.0M
Reka AI54 tok/s1.72 s$0.150$0.500262kfp8
Fireworks49 tok/s1.39 s$0.150$0.5001.0M
Cloudflare44 tok/s1.68 s$0.150$0.5001.3M
Z.ai39 tok/s5.10 s$0.075$0.2501.0Mfp8
NovitaAI38 tok/s3.37 s$0.132$0.4401.0Mfp8
Phala34 tok/s3.43 s$0.150$0.5001.0Mfp8
SiliconFlow34 tok/s2.35 s$0.150$0.5001.0Mfp8
StreamLake32 tok/s4.31 s$0.135$0.4501.0Mfp8
NextBit31 tok/s1.78 s$0.150$0.5001.0Mfp8
Sail Research27 tok/s1.38 s$0.150$0.5001.0Mfp8
io.net25 tok/s1.77 s$0.150$0.500262kfp8
DigitalOcean19 tok/s2.78 s$0.150$0.5001.0M
GMICloud18 tok/s3.35 s$0.113$0.3751.0Mfp8
Venice13 tok/s3.92 s$0.150$0.5001.0M
DeepInfra12 tok/s2.34 s$0.075$0.2501.0Mfp4
Wafer10 tok/s2.13 s$0.100$0.3501.0M

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.015 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0001$0.000134.5 s
Summarise a 30-page report12,000 / 600$0.0011$0.000538.7 s
Code edit6,000 / 1,500$0.0008$0.000651.2 s
Agentic coding session60,000 / 4,000$0.0055$0.00281.4 min
Structured extraction2,000 / 200$0.0002$0.000133.1 s

See also

Data as of 9 Sept 2026. Compare these configurations.