BenchLeader

Grok 4.5 vs Kimi K2.6

Verdict
  • Grok 4.5 leads on quality: 63.5 vs 60.7.
  • Grok 4.5 is stronger in agents & tools, human preference, knowledge, multimodal, reasoning.
  • Kimi K2.6 is stronger in coding, composite, long context, instruction following, maths.
  • Kimi K2.6 is 1.8× cheaper ($1.71 vs $3.00 per 1M blended).
MetricGrok 4.5Kimi K2.6
BenchLeader Index63.560.7
Agents & tools score58.645.1
Coding score59.860.4
Composite score66.268.6
Human preference score67.260.8
Knowledge score76.258.6
Long context score66.267.1
Multimodal score64.963.7
Reasoning score69.266.0
Instruction following score73.7
Maths score53.8
Blended price $/M$3.00$1.71
Output speed52 tok/s44 tok/s
Time to first answer13.6 s103.2 s
Context window500k262k
GPQA Diamond90.8%
FrontierMath Tiers 1–357.2%
FrontierMath Tier 425.6%
OTIS Mock AIME96.1%
SWE-bench Verified (Epoch)76.7%
SimpleQA Verified34.9%
SimpleBench70.0%
OSWorld-Verified 2.04.6%
SciCode53.5%
WeirdML46.4%55.9%
APEX-Agents34.2%18.9%
FrontierCode42.4%
ProofBench16.0%
Epoch Capabilities Index153.9151.0
LMArena Text14711461
LMArena Hard Prompts14951485
LMArena Coding15231514
LMArena WebDev15561509
LMArena Vision12911281
LMArena Agent3.9
LiveBench75.8%
LiveBench Reasoning87.2%
LiveBench Coding68.6%
LiveBench Agentic Coding56.5%
LiveBench Mathematics90.8%
LiveBench Data Analysis73.0%
LiveBench Language82.8%
AA Intelligence Index39.131.3
IFBench76.0%
AA-LCR79.3%81.0%
MMMU-Pro80.4%79.4%
AA-Omniscience25.35.3
Terminal-Bench Hard43.9%
GPQA Diamond (AA)93.1%91.1%
Humanity's Last Exam (AA)42.7%37.5%
SciCode (AA)55.0%51.5%
τ²-Bench Telecom (AA)95.9%
LiveCodeBench86.8%
MMLU-Pro87.6%
LegalBench84.7%
CorpFin66.7%
TaxEval74.7%
Terminal-Bench 2.1 (Vals)53.6%
SWE-bench (Vals)76.2%
GPQA Diamond (Vals)89.1%
Vals Index43.5
HiL-Bench18.7%
EQ-Bench 41202
Kagi LLM Benchmark83.5%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Grok 4.5 vs Kimi K2.6: questions

Is Grok 4.5 better than Kimi K2.6?
Grok 4.5 leads on quality: 63.5 vs 60.7. The BenchLeader Index combines every independent quality benchmark; Grok 4.5 is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Grok 4.5 better than Kimi K2.6 for coding?
Kimi K2.6 scores higher in coding (60 vs 60 on the category index, where 50 is average).
Is Grok 4.5 better than Kimi K2.6 for agentic tasks?
Grok 4.5 scores higher in agentic tasks (59 vs 45 on the category index, where 50 is average).
Which is cheaper, Grok 4.5 or Kimi K2.6?
Kimi K2.6 is cheaper: $1.71 against $3.00 per million tokens, blended at three input tokens per output token.
Which is faster, Grok 4.5 or Kimi K2.6?
Grok 4.5 streams faster: 52 against 44 output tokens per second.
Which has the larger context window?
Grok 4.5 accepts more context: 500k against 262k tokens.