BenchLeader

Grok 4.3 vs Kimi K2.6

Verdict
  • Grok 4.3 (medium) and Kimi K2.6 are level on quality (60.7 vs 60.7).
  • Grok 4.3 (medium) is stronger in agents & tools, instruction following, knowledge.
  • Kimi K2.6 is stronger in composite, long context, multimodal, coding, human preference, maths, reasoning.
  • They cost about the same ($1.56 per 1M blended).
  • Grok 4.3 (medium) streams 2.5× faster (112 vs 44 tokens per second).
MetricGrok 4.3 (medium)Kimi K2.6
BenchLeader Index60.760.7
Agents & tools score60.145.1
Composite score60.368.6
Instruction following score80.073.7
Knowledge score72.158.6
Long context score64.067.1
Multimodal score60.263.7
Coding score60.4
Human preference score60.8
Maths score53.8
Reasoning score66.0
Blended price $/M$1.56$1.71
Output speed112 tok/s44 tok/s
Time to first answer12.1 s103.2 s
Context window1M262k
GPQA Diamond90.8%
FrontierMath Tiers 1–357.2%
FrontierMath Tier 425.6%
OTIS Mock AIME96.1%
SWE-bench Verified (Epoch)76.7%
SimpleQA Verified34.9%
OSWorld-Verified 2.04.6%
SciCode53.5%
WeirdML55.9%
APEX-Agents18.9%
ProofBench16.0%
Epoch Capabilities Index151.0
LMArena Text1461
LMArena Hard Prompts1485
LMArena Coding1514
LMArena WebDev1509
LMArena Vision1281
AA Intelligence Index24.831.3
IFBench83.3%76.0%
AA-LCR75.0%81.0%
MMMU-Pro75.8%79.4%
AA-Omniscience16.75.3
Terminal-Bench Hard30.3%43.9%
GPQA Diamond (AA)89.0%91.1%
Humanity's Last Exam (AA)30.0%37.5%
SciCode (AA)51.5%
τ²-Bench Telecom (AA)91.2%95.9%
LiveCodeBench86.8%
MMLU-Pro87.6%
LegalBench84.7%
CorpFin66.7%
TaxEval74.7%
Terminal-Bench 2.1 (Vals)53.6%
SWE-bench (Vals)76.2%
GPQA Diamond (Vals)89.1%
Vals Index43.5
HiL-Bench18.7%
EQ-Bench 41202

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Grok 4.3 vs Kimi K2.6: questions

Is Grok 4.3 better than Kimi K2.6?
Grok 4.3 (medium) and Kimi K2.6 are level on quality (60.7 vs 60.7). The BenchLeader Index combines every independent quality benchmark; Kimi K2.6 is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Grok 4.3 better than Kimi K2.6 for agentic tasks?
Grok 4.3 scores higher in agentic tasks (60 vs 45 on the category index, where 50 is average).
Which is cheaper, Grok 4.3 or Kimi K2.6?
Grok 4.3 is cheaper: $1.56 against $1.71 per million tokens, blended at three input tokens per output token.
Which is faster, Grok 4.3 or Kimi K2.6?
Grok 4.3 streams faster: 112 against 44 output tokens per second.
Which has the larger context window?
Grok 4.3 accepts more context: 1M against 262k tokens.