BenchLeader

GLM 5.3 Flash vs Grok 4.6

Verdict
  • Grok 4.6 (medium) leads on quality: 63.9 vs 61.1.
  • GLM 5.3 Flash is stronger in agents & tools, coding, human preference, multimodal, reasoning.
  • Grok 4.6 (medium) is stronger in composite, knowledge, long context.
  • GLM 5.3 Flash is 25× cheaper ($0.119 vs $3.00 per 1M blended).
  • GLM 5.3 Flash streams 1.7× faster (90 vs 53 tokens per second).
MetricGLM 5.3 FlashGrok 4.6 (medium)
BenchLeader Index61.163.9
Agents & tools score52.1
Coding score64.364.1
Composite score61.783.4
Human preference score67.6
Knowledge score67.777.5
Long context score66.667.1
Multimodal score65.5
Reasoning score67.562.1
Blended price $/M$0.119$3.00
Output speed90 tok/s53 tok/s
Time to first answer24.9 s33.4 s
Context window1M500k
SciCode46.1%54.6%
Epoch Capabilities Index151.4
LMArena Text1474
LMArena Hard Prompts1496
LMArena Coding1534
LMArena WebDev1605
LMArena Vision1296
LMArena Agent2
LiveBench71.6%
LiveBench Reasoning77.6%
LiveBench Coding79.0%
LiveBench Agentic Coding56.8%
LiveBench Mathematics81.2%
LiveBench Data Analysis76.4%
LiveBench Language77.3%
AA Intelligence Index41.943.0
AA-LCR80.0%81.0%
AA-Omniscience7.528
GPQA Diamond (AA)91.2%93.5%
Humanity's Last Exam (AA)39.9%42.1%
SciCode (AA)51.6%55.9%
ARC-AGI-187.5%
ARC-AGI-261.3%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

GLM 5.3 Flash vs Grok 4.6: questions

Is GLM 5.3 Flash better than Grok 4.6?
Grok 4.6 (medium) leads on quality: 63.9 vs 61.1. The BenchLeader Index combines every independent quality benchmark; Grok 4.6 (medium) is ahead overall as of 2026-09-10, but check the category scores for your use.
Is GLM 5.3 Flash better than Grok 4.6 for coding?
GLM 5.3 Flash scores higher in coding (64 vs 64 on the category index, where 50 is average).
Which is cheaper, GLM 5.3 Flash or Grok 4.6?
GLM 5.3 Flash is cheaper: $0.119 against $3.00 per million tokens, blended at three input tokens per output token.
Which is faster, GLM 5.3 Flash or Grok 4.6?
GLM 5.3 Flash streams faster: 90 against 53 output tokens per second.
Which has the larger context window?
GLM 5.3 Flash accepts more context: 1M against 500k tokens.