BenchLeader

Claude Sonnet 4.6 vs Kimi K3

Verdict
  • Kimi K3 leads on quality: 66.5 vs 60.4.
  • Claude Sonnet 4.6 is stronger in instruction following, reasoning.
  • Kimi K3 is stronger in agents & tools, coding, composite, human preference, knowledge, long context, maths, multimodal.
  • They cost about the same ($6.00 per 1M blended).
MetricClaude Sonnet 4.6Kimi K3
BenchLeader Index60.466.5
Agents & tools score54.667.4
Coding score57.165.0
Composite score67.574.2
Human preference score62.071.2
Instruction following score56.9
Knowledge score62.368.2
Long context score66.671.1
Maths score62.778.1
Multimodal score60.665.0
Reasoning score65.2
Blended price $/M$6.00$6.00
Output speed42 tok/s38 tok/s
Time to first answer2.2 s56.3 s
Context window1M1.0M
GPQA Diamond87.4%
OTIS Mock AIME85.8%
SWE-bench Verified (Epoch)75.2%
Terminal-Bench53.4%
SciCode58.7%
APEX-Agents39.3%
FrontierCode24.3%44.2%
ProofBench87.0%
Epoch Capabilities Index152.3157.6
LMArena Text1472
LMArena Hard Prompts1504
LMArena Coding1528
LMArena WebDev1521
LMArena Vision1282
LMArena Agent-1
LiveBench79.2%
LiveBench Reasoning90.7%
LiveBench Coding81.5%
LiveBench Agentic Coding62.2%
LiveBench Mathematics84.4%
LiveBench Data Analysis78.7%
LiveBench Language85.5%
AA Intelligence Index30.443.8
IFBench56.6%
AA-LCR80.0%88.7%
MMMU-Pro73.3%80.5%
AA-Omniscience12.219.7
Terminal-Bench Hard53.0%
GPQA Diamond (AA)87.5%93.5%
Humanity's Last Exam (AA)33.6%46.9%
SciCode (AA)50.1%59.5%
τ²-Bench Telecom (AA)79.5%
AIME (Vals)92.3%
LiveCodeBench82.1%
MMLU-Pro87.3%
LegalBench82.1%86.0%
CorpFin65.3%71.6%
TaxEval77.1%75.7%
MedQA92.1%
Terminal-Bench 2.1 (Vals)57.3%80.9%
SWE-bench (Vals)77.4%93.4%
GPQA Diamond (Vals)85.6%92.9%
Vals Index50.657.8
MCP Atlas69.5%
EQ-Bench 412071339

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Claude Sonnet 4.6 vs Kimi K3: questions

Is Claude Sonnet 4.6 better than Kimi K3?
Kimi K3 leads on quality: 66.5 vs 60.4. The BenchLeader Index combines every independent quality benchmark; Kimi K3 is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Claude Sonnet 4.6 better than Kimi K3 for coding?
Kimi K3 scores higher in coding (65 vs 57 on the category index, where 50 is average).
Is Claude Sonnet 4.6 better than Kimi K3 for agentic tasks?
Kimi K3 scores higher in agentic tasks (67 vs 55 on the category index, where 50 is average).
Which is cheaper, Claude Sonnet 4.6 or Kimi K3?
Claude Sonnet 4.6 is cheaper: $6.00 against $6.00 per million tokens, blended at three input tokens per output token.
Which is faster, Claude Sonnet 4.6 or Kimi K3?
Claude Sonnet 4.6 streams faster: 42 against 38 output tokens per second.
Which has the larger context window?
Kimi K3 accepts more context: 1.0M against 1M tokens.