BenchLeader

Kimi K3 vs Qwen3.8 Max

Verdict
  • Kimi K3 (max) leads on quality: 65.5 vs 64.2.
  • Kimi K3 (max) is stronger in agents & tools, coding, composite, human preference, knowledge, long context.
  • Qwen3.8 Max (max) is stronger in maths, multimodal, reasoning.
  • Qwen3.8 Max (max) is 2.0× cheaper ($3.00 vs $6.00 per 1M blended).
MetricKimi K3 (max)Qwen3.8 Max (max)
BenchLeader Index65.564.2
Agents & tools score66.258.2
Coding score64.764.6
Composite score79.070.7
Human preference score68.668.0
Knowledge score63.262.7
Long context score69.765.4
Maths score63.865.1
Multimodal score63.766.2
Reasoning score66.672.2
Blended price $/M$6.00$3.00
Output speed41 tok/s37 tok/s
Time to first answer52.2 s58.9 s
Context window1.0M1M
GPQA Diamond93.1%–
FrontierMath Tiers 1–372.2%–
FrontierMath Tier 439.0%–
OTIS Mock AIME97.2%–
SimpleQA Verified50.6%–
Terminal-Bench–27.0%
SimpleBench60.7%–
SciCode59.5%53.2%
WeirdML82.6%–
APEX-Agents–63.3%
ProofBench–58.0%
Epoch Capabilities Index–156.4
LMArena Text14881483
LMArena Hard Prompts15171504
LMArena Coding15401524
LMArena WebDev16541672
LMArena Vision–1314
LMArena Agent3.82.3
LiveBench–78.5%
LiveBench Reasoning–88.2%
LiveBench Coding–72.9%
LiveBench Agentic Coding–64.7%
LiveBench Mathematics–91.3%
LiveBench Data Analysis–78.4%
LiveBench Language–79.7%
LiveBench Instruction Following–74.1%
AA Intelligence Index v4.3.243.645.4
AA-LCR88.7%80.3%
MMMU-Pro80.5%82.8%
AA-Omniscience19.712.0
GPQA Diamond (AA)93.5%92.8%
Humanity's Last Exam (AA)46.9%43.1%
SciCode (AA)59.5%53.2%
LiveCodeBench87.2%87.8%
MMLU-Pro88.0%88.6%
IOI48.9%68.9%
LegalBench–83.6%
CorpFin–65.8%
TaxEval–75.5%
Terminal-Bench 2.1 (Vals)–67.4%
SWE-bench (Vals)–85.6%
GPQA Diamond (Vals)–93.7%
Vals Index–48.3
MCP Atlas82.3%–
ARC-AGI-194.5%–
ARC-AGI-260.4%–
CritPt23.4%20.0%
GDPval-AA v2.151.7%58.6%
τ³-Banking (AA)46.0%51.3%
ITBench SRE (AA)47.7%40.3%
Analyst Agent (AA)38.8%45.0%
APEX-Agents (AA)41.3%42.4%
BioMysteryBench71.5%–
Code Migration16.1%24.0%
CyberBench75.2%28.6%
Excel Modeling Benchmark66.4%60.1%
Finance Agent v2–50.6%
Harvey's Legal Agent Benchmark10.8%10.4%
Legal Research Bench–47.6%
MedCode–40.7%
MedScribe–85.0%
MMMU-Pro (Vals)–88.0%
MortgageTax–64.0%
MysteryMechanism–23.9%
ProgramBench–0.0%
Public Benefits Bench68.3%67.1%
SAGE–51.3%
SkillsBench–42.0%
Tax Agent Bench68.7%66.0%
Terminal-Bench 4.0 (Vals)19.7%34.3%
Terminal-Bench Science2.9%1.4%
Time Horizon Index: KSP10.5%–
Vals Multimodal Index–65.4%
Vibe Code Bench 1-10018.2%12.8%
Vibe Code Bench v1.1–64.7%
FORTRESS26.6%–
LMArena Maths14981497
LMArena Creative Writing14601470
LMArena Instruction Following14881474
LMArena Multi-turn15001492
LMArena Longer Queries15041492
Chess Puzzles39.0%–
Mystery Game Puzzles26.0%–
Surface Evolver Bench93.0%–
DeepSWE v1.168.5%–
LMCA52.7%–
DTBench91.2%–
ForecastBench61.1%–
ALE-Bench1524.5–
GDP.pdf19.0%–
FrontierSWE25.9%–
Terminal-Bench 4.0 (AA)12.6%38.9%
Terminal-Bench 2.1 (AA)85.0%88.8%
AutomationBench58.3%56.2%
GDP.pdf22.0%22.8%
MLCR38.3%20.0%
Harvey LAB5.3%–
EnterpriseOps-Gym45.3%47.6%
AA-Omniscience: accuracy47.6%31.9%
AA-Omniscience: non-hallucination46.8%71.2%
AA-Briefcase v1.115011617
AA Openness Index38.9–

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

Kimi K3 vs Qwen3.8 Max: questions

Is Kimi K3 better than Qwen3.8 Max?
Kimi K3 (max) leads on quality: 65.5 vs 64.2. The BenchLeader Index combines every independent quality benchmark; Kimi K3 (max) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is Kimi K3 better than Qwen3.8 Max for coding?
Kimi K3 scores higher in coding (65 vs 65 on the category index, where 50 is average).
Is Kimi K3 better than Qwen3.8 Max for agentic tasks?
Kimi K3 scores higher in agentic tasks (66 vs 58 on the category index, where 50 is average).
Which is cheaper, Kimi K3 or Qwen3.8 Max?
Qwen3.8 Max is cheaper: $3.00 against $6.00 per million tokens, blended at three input tokens per output token.
Which is faster, Kimi K3 or Qwen3.8 Max?
Kimi K3 streams faster: 41 against 37 output tokens per second.
Which has the larger context window?
Kimi K3 accepts more context: 1.0M against 1M tokens.