BenchLeader

Kimi K3

Best configuration ranks #18 of 610 on the BenchLeader Index at 66.5 ±3.4. Last measured 5 Sept 2026. Released 16 Jul 2026.

Blended price
$6.00/M
$3.00 in · $15.00 out
Output speed
37 tok/s
First answer
58 s
first token 1.84 s
Context
1.0M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index67
  2. Coding65
  3. Agents & tools68
  4. Maths78
  5. Knowledge68
  6. Human preference71
  7. Multimodal65
  8. Long context71
  9. Composite74

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
low58.8#11737 tok/s58 s$0.0057Coding 59 · Composite 72 · Knowledge 66 · Long context 66 · Maths 54 · Multimodal 63 · Reasoning 51
high37 tok/s58 sMaths 65 · Reasoning 63
max62.7#5137 tok/s58 sAgents & tools 63 · Coding 70 · Human preference 69 · Knowledge 60 · Maths 63 · Reasoning 64
thinking37 tok/s58 sMaths 65
defaultbest66.5#1837 tok/s58 s$0.0057Agents & tools 68 · Coding 65 · Composite 74 · Human preference 71 · Knowledge 68 · Long context 71 · Maths 78 · Multimodal 65

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarklowhighmaxthinkingdefaultSourceTrend
GPQA Diamond84.8%#8191.9%#2293.1%#17Epoch AI Benchmarking Hub
SimpleBench60.7%#28SimpleBench
LMArena Hard Prompts1519#7LMArena
LiveBench Reasoningnot in index90.7%#5LiveBench
GPQA Diamond (AA)not in index84.2%#14493.5%#14Artificial Analysis
Humanity's Last Exam (AA)not in index25.0%#13546.9%#24Artificial Analysis
GPQA Diamond (Vals)not in index92.9%#12Vals AI
ARC-AGI-165.7%#10186.7%#6694.5%#24ARC Prize
ARC-AGI-212.4%#10355.0%#6660.4%#56ARC Prize

Coding

BenchmarklowhighmaxthinkingdefaultSourceTrend
SciCode51.2%#5358.7%#558.7%#5SciCode
WeirdML82.6%#15WeirdML
FrontierCode44.2%#8Cognition
LMArena Coding1542#6LMArena
LMArena WebDev1674#5LMArena
LiveBench Codingnot in index81.5%#9LiveBench
SciCode (AA)not in index52.7%#5159.5%#6Artificial Analysis
LiveCodeBench87.2%#17Vals AI
SWE-bench (Vals)not in index93.4%#8Vals AI

Agents & tools

BenchmarklowhighmaxthinkingdefaultSourceTrend
APEX-Agents39.3%#11Mercor
LMArena Agent6.6#8LMArena
LiveBench Agentic Codingnot in index62.2%#6LiveBench
Terminal-Bench 2.1 (Vals)80.9%#6Vals AI
MCP Atlas82.3%#6Scale AI SEAL

Maths

BenchmarklowhighmaxthinkingdefaultSourceTrend
FrontierMath Tiers 1–372.2%#15Epoch AI Benchmarking Hub
FrontierMath Tier 439.0%#22Epoch AI Benchmarking Hub
OTIS Mock AIME68.9%#12693.3%#4497.2%#23Epoch AI Benchmarking Hub
ProofBench87.0%#5Vals AI
LiveBench Mathematicsnot in index84.4%#41LiveBench
AIME 202696.7%#6MathArena
HMMT February 202697.0%#3MathArena
MathArena Apex65.6%#4MathArena

Knowledge

BenchmarklowhighmaxthinkingdefaultSourceTrend
SimpleQA Verified50.6%#20Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index78.7%#14LiveBench
AA-Omniscience3.9#7719.7#43Artificial Analysis
MMLU-Pro88.0%#22Vals AI
LegalBench86.0%#14Vals AI
CorpFin71.6%#3Vals AI
TaxEval75.7%#12Vals AI

Instruction following

BenchmarklowhighmaxthinkingdefaultSourceTrend
LiveBench Languagenot in index85.5%#8LiveBench

Human preference

BenchmarklowhighmaxthinkingdefaultSourceTrend
LMArena Text1489#12LMArena
EQ-Bench 41339#3EQ-Bench

Multimodal

BenchmarklowhighmaxthinkingdefaultSourceTrend
MMMU-Pro78.5%#5280.5%#31Artificial Analysis

Long context

BenchmarklowhighmaxthinkingdefaultSourceTrend
AA-LCR79.3%#6588.7%#1Artificial Analysis

Composite

BenchmarklowhighmaxthinkingdefaultSourceTrend
Epoch Capabilities Indexnot in index157.6#11Epoch AI Benchmarking Hub
LiveBench79.2%#8LiveBench
AA Intelligence Index34.5#5743.8#24Artificial Analysis
Vals Indexnot in index57.8#15Vals AI

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Modal80 tok/s0.99 s$3.00$15.001.0Mmxfp4
Relace68 tok/s1.52 s$2.40$12.001.0Mfp4
Together68 tok/s1.28 s$3.00$15.001.0M
Sail Research57 tok/s1.68 s$2.60$13.001.0Mfp4
Parasail55 tok/s1.37 s$3.00$15.001.0Mfp4
Makora54 tok/s1.43 s$2.55$12.751.0M
Wafer50 tok/s0.74 s$3.00$12.751.0M
Fireworks (US)49 tok/s2.24 s$3.30$16.501.0M
Fireworks44 tok/s2.16 s$3.00$15.001.0M
Fireworks Fast40 tok/s1.13 s$4.50$22.501.0M
Baseten36 tok/s1.80 s$3.00$15.001.0Mfp8
Moonshot AI26 tok/s3.90 s$3.00$15.001.0Mmxfp4
Alibaba Cloud Int.25 tok/s1.84 s$3.45$17.251.0M
DigitalOcean24 tok/s2.52 s$2.85$14.251.0M
Phala22 tok/s4.12 s$3.00$15.001.0M
Chutes21 tok/s2.54 s$3.00$15.001.0Mmxfp4
Morph Fast13 tok/s3.24 s$6.00$22.501.0Mfp4
Morph12 tok/s3.45 s$2.80$14.001.0Mfp4
DeepInfra12 tok/s3.70 s$2.85$14.251.0Mbf16

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$4.13$8.25$12.38$16.50Jul 26Aug 26
input outputnow $2.90 in · $14.00 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.300 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0057$0.00491.1 min
Summarise a 30-page report12,000 / 600$0.045$0.0211.2 min
Code edit6,000 / 1,500$0.041$0.0281.6 min
Agentic coding session60,000 / 4,000$0.240$0.1182.8 min
Structured extraction2,000 / 200$0.0090$0.00501.1 min

See also

Data as of 9 Sept 2026. Compare these configurations.