BenchLeader

Kimi K3 vs Qwen3.5 397B A17B

Verdict
  • Kimi K3 leads on quality: 66.5 vs 58.7.
  • Kimi K3 is stronger in coding, composite, human preference, knowledge, long context, maths, multimodal.
  • Qwen3.5 397B A17B is stronger in agents & tools, instruction following, reasoning.
  • Qwen3.5 397B A17B is 16× cheaper ($0.387 vs $6.00 per 1M blended).
  • Qwen3.5 397B A17B streams 2.1× faster (78 vs 38 tokens per second).
MetricKimi K3Qwen3.5 397B A17B
BenchLeader Index66.558.7
Agents & tools score67.469.4
Coding score65.051.7
Composite score74.253.1
Human preference score71.263.6
Knowledge score68.249.5
Long context score71.165.2
Maths score78.150.1
Multimodal score65.061.6
Instruction following score76.1
Reasoning score63.0
Blended price $/M$6.00$0.387
Output speed38 tok/s78 tok/s
Time to first answer56.3 s42.7 s
Context window1.0M262k
GPQA Diamond85.9%
FrontierMath Tiers 1–329.5%
OTIS Mock AIME88.9%
SciCode58.7%
APEX-Agents39.3%
FrontierCode44.2%
ProofBench87.0%
Epoch Capabilities Index157.6147.0
LMArena Text1441
LMArena Hard Prompts1463
LMArena Coding1491
LMArena WebDev1399
LMArena Vision1265
LiveBench79.2%
LiveBench Reasoning90.7%
LiveBench Coding81.5%
LiveBench Agentic Coding62.2%
LiveBench Mathematics84.4%
LiveBench Data Analysis78.7%
LiveBench Language85.5%
AA Intelligence Index43.819.1
IFBench78.8%
AA-LCR88.7%77.3%
MMMU-Pro80.5%77.3%
AA-Omniscience19.7-30.8
Terminal-Bench Hard40.9%
GPQA Diamond (AA)93.5%89.3%
Humanity's Last Exam (AA)46.9%29.0%
SciCode (AA)59.5%44.8%
τ²-Bench Telecom (AA)95.6%
LegalBench86.0%
CorpFin71.6%
TaxEval75.7%
Terminal-Bench 2.1 (Vals)80.9%
SWE-bench (Vals)93.4%
GPQA Diamond (Vals)92.9%
Vals Index57.8
AIME 202694.2%
HMMT February 202687.9%
EQ-Bench 41339

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Kimi K3 vs Qwen3.5 397B A17B: questions

Is Kimi K3 better than Qwen3.5 397B A17B?
Kimi K3 leads on quality: 66.5 vs 58.7. The BenchLeader Index combines every independent quality benchmark; Kimi K3 is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Kimi K3 better than Qwen3.5 397B A17B for coding?
Kimi K3 scores higher in coding (65 vs 52 on the category index, where 50 is average).
Is Kimi K3 better than Qwen3.5 397B A17B for agentic tasks?
Qwen3.5 397B A17B scores higher in agentic tasks (69 vs 67 on the category index, where 50 is average).
Which is cheaper, Kimi K3 or Qwen3.5 397B A17B?
Qwen3.5 397B A17B is cheaper: $0.387 against $6.00 per million tokens, blended at three input tokens per output token.
Which is faster, Kimi K3 or Qwen3.5 397B A17B?
Qwen3.5 397B A17B streams faster: 78 against 38 output tokens per second.
Which has the larger context window?
Kimi K3 accepts more context: 1.0M against 262k tokens.