BenchLeader

Kimi K2.5

Best configuration ranks #162 of 610 on the BenchLeader Index at 56.1 ±3.0. Last measured 13 Feb 2026stale: no new result in six months. Released 27 Jan 2026.

Blended price
$1.20/M
$0.600 in · $3.00 out
Output speed
37 tok/s
First answer
1.64 s
Context
262k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index56
  2. Reasoning53
  3. Coding53
  4. Agents & tools48
  5. Maths65
  6. Knowledge52
  7. Instruction following62
  8. Multimodal55
  9. Long context65
  10. Composite59

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning52.7#24137 tok/s1.64 s$0.0011Agents & tools 50 · Composite 54 · Instruction following 46 · Knowledge 58 · Long context 60 · Multimodal 58
high37 tok/s1.64 s$0.0010Coding 62
thinking55.3#17837 tok/s1.64 s$0.0010Agents & tools 32 · Coding 55 · Human preference 64 · Knowledge 60 · Maths 53 · Multimodal 62 · Reasoning 68
defaultbest56.1#16237 tok/s1.64 s$0.0011Agents & tools 48 · Coding 53 · Composite 59 · Instruction following 62 · Knowledge 52 · Long context 65 · Maths 65 · Multimodal 55 · Reasoning 53

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoninghighthinkingdefaultSourceTrend
GPQA Diamond87.6%#60Epoch AI Benchmarking Hub
Humanity's Last Exam24.4%#14Scale AI / CAIS
SimpleBench46.8%#48SimpleBench
LMArena Hard Prompts1472#68LMArena
GPQA Diamond (AA)not in index78.9%#20087.9%#92Artificial Analysis
Humanity's Last Exam (AA)not in index13.2%#21030.7%#99Artificial Analysis
GPQA Diamond (Vals)not in index84.1%#52Vals AI
Kagi LLM Benchmark78.5%#1063.8%#40Kagi LLM Benchmark
ARC-AGI-165.3%#103ARC Prize
ARC-AGI-211.8%#104ARC Prize

Coding

Benchmarkno reasoninghighthinkingdefaultSourceTrend
SWE-bench Verified (Epoch)73.8%#17Epoch AI Benchmarking Hub
SciCode49.0%#66SciCode
WeirdML45.6%#81WeirdML
LMArena Coding1502#60LMArena
LMArena WebDev1436#59LMArena
LiveCodeBench83.9%#44Vals AI
IOI17.7%#26Vals AI
SWE-bench (Vals)not in index70.0%#58Vals AI
SWE-bench Verified (bash only)70.8%#11SWE-bench
SWE-bench Verified (any scaffold)not in index70.8%#15SWE-bench

Agents & tools

Benchmarkno reasoninghighthinkingdefaultSourceTrend
Terminal-Bench43.2%#34Terminal-Bench
APEX-Agents14.4%#47Mercor
Terminal-Bench Hard18.9%#16134.9%#80Artificial Analysis
τ²-Bench Telecom (AA)not in index81.3%#11495.9%#16Artificial Analysis
Terminal-Bench 2.1 (Vals)42.0%#54Vals AI
MCP Atlas64.4%#22Scale AI SEAL

Maths

Benchmarkno reasoninghighthinkingdefaultSourceTrend
OTIS Mock AIME92.2%#52Epoch AI Benchmarking Hub
AIME (Vals)95.6%#8Vals AI
AIME 202695.8%#13MathArena
HMMT February 202687.1%#19MathArena
MathArena Apex8.8%#23MathArena

Knowledge

Benchmarkno reasoninghighthinkingdefaultSourceTrend
SimpleQA Verified34.3%#43Epoch AI Benchmarking Hub
AA-Omniscience-13.8#158-7.3#124Artificial Analysis
MMLU-Pro85.9%#50Vals AI
CorpFin68.3%#9Vals AI
TaxEval74.2%#39Vals AI
MedQA94.4%#17Vals AI
PRBench Finance46.5%#14Scale AI SEAL
PRBench Legal43.8%#18Scale AI SEAL
MultiNRC35.2%#21Scale AI SEAL

Instruction following

Benchmarkno reasoninghighthinkingdefaultSourceTrend
IFBench43.7%#21670.2%#69Artificial Analysis
MultiChallenge61.4%#9Scale AI SEAL
TutorBench54.6%#6Scale AI SEAL

Human preference

Benchmarkno reasoninghighthinkingdefaultSourceTrend
LMArena Text1451#65LMArena

Multimodal

Benchmarkno reasoninghighthinkingdefaultSourceTrend
LMArena Vision1267#36LMArena
MMMU-Pro73.1%#9975.4%#74Artificial Analysis
VISTA41.9%#37Scale AI SEAL

Long context

Benchmarkno reasoninghighthinkingdefaultSourceTrend
Fiction.LiveBench 120k78.1%#6Fiction.live
AA-LCR67.3%#18278.0%#82Artificial Analysis

Composite

Benchmarkno reasoninghighthinkingdefaultSourceTrend
Epoch Capabilities Indexnot in index148.0#51Epoch AI Benchmarking Hub
AA Intelligence Index19.4#18023.5#131Artificial Analysis
Vals Indexnot in index26.3#44Vals AI

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Venice104 tok/s1.33 s$0.532$3.33256k
Amazon Bedrock44 tok/s1.74 s$0.600$3.00262k
AtlasCloud37 tok/s0.95 s$0.490$2.50262kint4
Phala37 tok/s1.83 s$0.600$3.00262k
NovitaAI36 tok/s1.54 s$0.570$2.85262k
SiliconFlow19 tok/s3.27 s$0.450$2.25262kint4

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$0.825$1.65$2.48$3.30Feb 26Mar 26Apr 26May 26Jun 26Jul 26Aug 26Sept 26
input outputnow $0.450 in · $2.25 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.100 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0011$0.00109.7 s
Summarise a 30-page report12,000 / 600$0.0090$0.004517.9 s
Code edit6,000 / 1,500$0.0081$0.005942.2 s
Agentic coding session60,000 / 4,000$0.048$0.0251.8 min
Structured extraction2,000 / 200$0.0018$0.00117.0 s

See also

Data as of 9 Sept 2026. Compare these configurations.