BenchLeader

Grok 4.5 vs Kimi K3

Verdict
  • Kimi K3 leads on quality: 66.5 vs 63.4.
  • Grok 4.5 is stronger in knowledge, reasoning.
  • Kimi K3 is stronger in agents & tools, coding, composite, human preference, long context, multimodal, maths.
  • Grok 4.5 is 2.0× cheaper ($3.00 vs $6.00 per 1M blended).
  • Grok 4.5 streams 1.5× faster (56 vs 37 tokens per second).
MetricGrok 4.5Kimi K3
BenchLeader Index63.466.5
Agents & tools score58.667.5
Coding score59.964.9
Composite score66.074.1
Human preference score66.971.2
Knowledge score76.068.1
Long context score66.271.1
Multimodal score64.865.1
Reasoning score68.9
Maths score78.1
Blended price $/M$3.00$6.00
Output speed56 tok/s37 tok/s
Time to first answer12.7 s57.7 s
Context window500k1.0M
SimpleBench70.0%
SciCode58.7%
WeirdML46.4%
APEX-Agents34.2%39.3%
FrontierCode42.4%44.2%
ProofBench87.0%
Epoch Capabilities Index153.9157.6
LMArena Text1471
LMArena Hard Prompts1495
LMArena Coding1523
LMArena WebDev1556
LMArena Vision1291
LMArena Agent3.9
LiveBench75.8%79.2%
LiveBench Reasoning87.2%90.7%
LiveBench Coding68.6%81.5%
LiveBench Agentic Coding56.5%62.2%
LiveBench Mathematics90.8%84.4%
LiveBench Data Analysis73.0%78.7%
LiveBench Language82.8%85.5%
AA Intelligence Index39.143.8
AA-LCR79.3%88.7%
MMMU-Pro80.4%80.5%
AA-Omniscience25.319.7
GPQA Diamond (AA)93.1%93.5%
Humanity's Last Exam (AA)42.7%46.9%
SciCode (AA)55.0%59.5%
LegalBench86.0%
CorpFin71.6%
TaxEval75.7%
Terminal-Bench 2.1 (Vals)80.9%
SWE-bench (Vals)93.4%
GPQA Diamond (Vals)92.9%
Vals Index57.8
EQ-Bench 41339
Kagi LLM Benchmark83.5%

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.