BenchLeader

Grok 4.20 vs Kimi K3

Verdict
  • Kimi K3 (max) leads on quality: 67.2 vs 59.9.
  • Grok 4.20 (thinking) is stronger in instruction following.
  • Kimi K3 (max) is stronger in agents & tools, coding, composite, human preference, knowledge, long context, maths, multimodal, reasoning.
  • Grok 4.20 (thinking) is 3.8× cheaper ($1.56 vs $6.00 per 1M blended).
  • Grok 4.20 (thinking) streams 2.7× faster (101 vs 37 tokens per second).
MetricGrok 4.20 (thinking)Kimi K3 (max)
BenchLeader Index59.967.2
Agents & tools score51.767.5
Coding score54.767.0
Composite score61.584.5
Human preference score67.168.8
Instruction following score79.9
Knowledge score56.864.5
Long context score60.771.0
Maths score55.665.4
Multimodal score59.965.1
Reasoning score60.768.6
Blended price $/M$1.56$6.00
Output speed101 tok/s37 tok/s
Time to first answer21.4 s58.3 s
Context window1M1.0M
GPQA Diamond89.3%93.1%
FrontierMath Tiers 1–344.9%72.2%
FrontierMath Tier 417.1%39.0%
OTIS Mock AIME92.2%97.2%
SimpleQA Verified30.2%50.6%
SimpleBench60.7%
SciCode58.7%
WeirdML82.6%
ProofBench14.0%
LMArena Text14721485
LMArena Hard Prompts14881514
LMArena Coding15101538
LMArena WebDev13741674
LMArena Vision1263
LMArena Agent6.2
AA Intelligence Index25.743.8
IFBench82.9%
AA-LCR69.0%88.7%
MMMU-Pro74.6%80.5%
AA-Omniscience14.819.7
Terminal-Bench Hard40.9%
GPQA Diamond (AA)91.1%93.5%
Humanity's Last Exam (AA)34.5%46.9%
SciCode (AA)59.5%
τ²-Bench Telecom (AA)96.5%
AIME (Vals)96.5%
LiveCodeBench84.3%87.2%
MMLU-Pro86.3%88.0%
IOI48.9%
LegalBench77.7%
CorpFin63.7%
TaxEval74.1%
MedQA94.5%
Terminal-Bench 2.1 (Vals)44.2%
SWE-bench (Vals)72.2%
GPQA Diamond (Vals)88.6%
Vals Index17.6
MCP Atlas82.3%
Kagi LLM Benchmark75.0%
ARC-AGI-189.5%94.5%
ARC-AGI-265.1%60.4%
ARC-AGI-30.1%
CritPt6.6%23.4%
GDPval (AA)52.4%
τ²-Bench Banking (AA)46.0%
ITBench SRE (AA)47.7%
Analyst Agent (AA)38.8%
APEX-Agents (AA)14.2%41.3%
BioMysteryBench71.5%
CaseLaw v254.5%
Code Migration0.3%16.1%
Excel Modeling Benchmark11.7%66.4%
Finance Agent v228.5%
Harvey's Legal Agent Benchmark0.0%
Legal Research Bench13.9%
MedCode32.2%
MedScribe63.4%
MMMU-Pro (Vals)83.5%
MortgageTax45.4%
SAGE38.2%
Tax Agent Bench68.7%
Terminal-Bench 2.0 (Vals)40.5%
Terminal-Bench 4.0 (Vals)12.6%
Terminal-Bench Science2.9%
Time Horizon Index: KSP10.5%
Vals Multimodal Index39.1%
Vibe Code Bench 1-10018.2%
Vibe Code Bench v1.14.1%
FORTRESS26.6%
LMArena Maths14671501
LMArena Creative Writing14461458
LMArena Instruction Following14451484
LMArena Multi-turn14801496
LMArena Longer Queries14651500
LMArena Document1439
Chess Puzzles24.0%39.0%
Mystery Game Puzzles26.0%
Surface Evolver Bench93.0%
DeepSWE68.5%
LMCA52.7%
DTBench91.2%
ForecastBench60.7%
CursorBench60.8%
ALE-Bench1524.5
GDP.pdf19.0%
FrontierSWE25.9%

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

Grok 4.20 vs Kimi K3: questions

Is Grok 4.20 better than Kimi K3?
Kimi K3 (max) leads on quality: 67.2 vs 59.9. The BenchLeader Index combines every independent quality benchmark; Kimi K3 (max) is ahead overall as of 2026-09-19, but check the category scores for your use.
Is Grok 4.20 better than Kimi K3 for coding?
Kimi K3 scores higher in coding (67 vs 55 on the category index, where 50 is average).
Is Grok 4.20 better than Kimi K3 for agentic tasks?
Kimi K3 scores higher in agentic tasks (68 vs 52 on the category index, where 50 is average).
Which is cheaper, Grok 4.20 or Kimi K3?
Grok 4.20 is cheaper: $1.56 against $6.00 per million tokens, blended at three input tokens per output token.
Which is faster, Grok 4.20 or Kimi K3?
Grok 4.20 streams faster: 101 against 37 output tokens per second.
Which has the larger context window?
Kimi K3 accepts more context: 1.0M against 1M tokens.