BenchLeader

GLM-4.6

Best configuration ranks #254 of 610 on the BenchLeader Index at 52.3 ±9.7 (thinking reasoning effort). Released 30 Sept 2025.

Blended price
$0.963/M
$0.550 in · $2.20 out
Output speed
44 tok/s
First answer
48 s
first token 2.74 s
Context
200k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index52
  2. Agents & tools72
  3. Knowledge44
  4. Instruction following45
  5. Long context53
  6. Composite52

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
thinkingbest52.3#25444 tok/s48 s$0.0009Agents & tools 72 · Composite 52 · Instruction following 45 · Knowledge 44 · Long context 53
default47.6#34965 tok/s2.74 s$0.0009Agents & tools 40 · Coding 44 · Composite 48 · Human preference 61 · Instruction following 40 · Knowledge 51 · Long context 38 · Maths 51 · Reasoning 53

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkthinkingdefaultSourceTrend
LMArena Hard Prompts1442#104LMArena
GPQA Diamond (AA)not in index78.0%#20963.2%#339Artificial Analysis
Humanity's Last Exam (AA)not in index14.5%#1985.5%#349Artificial Analysis
GPQA Diamond (Vals)not in index74.5%#80Vals AI
Kagi LLM Benchmark47.4%#90Kagi LLM Benchmark

Coding

BenchmarkthinkingdefaultSourceTrend
SciCode38.4%#120SciCode
LMArena Coding1458#118LMArena
LMArena WebDev1341#93LMArena
LiveCodeBench81.0%#62Vals AI
IOI4.3%#46Vals AI
SWE-Bench Pro9.7%#22Scale AI SEAL
SWE-bench Verified (bash only)55.4%#28SWE-bench
SWE-bench Verified (any scaffold)not in index55.4%#36SWE-bench

Agents & tools

Maths

BenchmarkthinkingdefaultSourceTrend
AIME (Vals)92.7%#17Vals AI
MGSM89.8%#42Vals AI
MathArena Apex0.5%#41MathArena

Knowledge

BenchmarkthinkingdefaultSourceTrend
AA-Omniscience-41.9#265-31.7#222Artificial Analysis
MMLU-Pro82.2%#76Vals AI
LegalBench79.6%#87Vals AI
CorpFin56.8%#86Vals AI
TaxEval66.2%#110Vals AI
MedQA92.2%#33Vals AI

Instruction following

BenchmarkthinkingdefaultSourceTrend
IFBench43.4%#21936.7%#282Artificial Analysis

Human preference

BenchmarkthinkingdefaultSourceTrend
LMArena Text1425#102LMArena

Long context

BenchmarkthinkingdefaultSourceTrend
AA-LCR54.0%#23926.3%#335Artificial Analysis

Composite

BenchmarkthinkingdefaultSourceTrend
Epoch Capabilities Indexnot in index140.8#95Epoch AI Benchmarking Hub
AA Intelligence Index18.5#19014.9#223Artificial Analysis

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Venice60 tok/s1.07 s$0.430$1.75198kfp4
AtlasCloud44 tok/s1.15 s$0.600$2.20203kfp8
NovitaAI28 tok/s2.35 s$0.550$2.20205kbf16
Z.ai27 tok/s12 s$0.600$2.20203kfp4
DeepInfra21 tok/s1.51 s$0.500$2.00203kfp4

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$0.605$1.21$1.82$2.42Jan 26Feb 26Mar 26Apr 26May 26Jun 26Jul 26Aug 26
input outputnow $0.600 in · $2.20 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.110 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0009$0.000754.8 s
Summarise a 30-page report12,000 / 600$0.0079$0.00401.0 min
Code edit6,000 / 1,500$0.0066$0.00461.4 min
Agentic coding session60,000 / 4,000$0.042$0.0222.3 min
Structured extraction2,000 / 200$0.0015$0.000952.6 s

See also

Data as of 9 Sept 2026. Compare these configurations.