BenchLeader

Kimi K3 vs Qwen3 8

Verdict
  • Kimi K3 leads on quality: 66.5 vs 63.9.
  • Kimi K3 is stronger in agents & tools, composite, human preference, knowledge, long context, maths.
  • Qwen3 8 (max) is stronger in coding, multimodal, reasoning.
  • Qwen3 8 (max) is 2.3× cheaper ($2.67 vs $6.00 per 1M blended).
MetricKimi K3Qwen3 8 (max)
BenchLeader Index66.563.9
Agents & tools score67.556.1
Coding score64.971.7
Composite score74.170.8
Human preference score71.267.9
Knowledge score68.162.3
Long context score71.165.7
Maths score78.1
Multimodal score65.167.2
Reasoning score67.7
Blended price $/M$6.00$2.67
Output speed37 tok/s39 tok/s
Time to first answer57.7 s53.2 s
Context window1.0M1M
SciCode58.7%
APEX-Agents39.3%
FrontierCode44.2%
ProofBench87.0%
Epoch Capabilities Index157.6156.6
LMArena Text1480
LMArena Hard Prompts1502
LMArena Coding1520
LMArena WebDev1685
LMArena Vision1313
LMArena Agent3.9
LiveBench79.2%78.5%
LiveBench Reasoning90.7%88.2%
LiveBench Coding81.5%72.9%
LiveBench Agentic Coding62.2%64.7%
LiveBench Mathematics84.4%91.3%
LiveBench Data Analysis78.7%78.4%
LiveBench Language85.5%79.7%
AA Intelligence Index43.840.3
AA-LCR88.7%78.3%
MMMU-Pro80.5%82.3%
AA-Omniscience19.73.4
GPQA Diamond (AA)93.5%92.7%
Humanity's Last Exam (AA)46.9%43.0%
SciCode (AA)59.5%53.2%
LiveCodeBench87.8%
MMLU-Pro88.6%
IOI73.0%
LegalBench86.0%83.6%
CorpFin71.6%65.8%
TaxEval75.7%75.5%
Terminal-Bench 2.1 (Vals)80.9%67.4%
SWE-bench (Vals)93.4%85.6%
GPQA Diamond (Vals)92.9%93.7%
Vals Index57.851.8
EQ-Bench 41339

Data as of 2026-09-09. Best configuration of each model; every score links to its source on the model pages.