BenchLeader

Kimi K2.6

Best configuration ranks #82 of 610 on the BenchLeader Index at 60.6 ±3.9. Last measured 8 Sept 2026. Released 20 Apr 2026.

Blended price
$1.71/M
$0.950 in · $4.00 out
Output speed
52 tok/s
First answer
89 s
first token 1.40 s
Context
262k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index61
  2. Reasoning66
  3. Coding60
  4. Agents & tools45
  5. Maths54
  6. Knowledge59
  7. Instruction following73
  8. Human preference61
  9. Multimodal64
  10. Long context67
  11. Composite68

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning55.3#18051 tok/s2.32 s$0.0016Agents & tools 67 · Composite 59 · Instruction following 46 · Knowledge 60 · Long context 61
thinking52 tok/s89 s$0.0015Composite 38 · Maths 57
defaultbest60.6#8252 tok/s89 s$0.0016Agents & tools 45 · Coding 60 · Composite 68 · Human preference 61 · Instruction following 73 · Knowledge 59 · Long context 67 · Maths 54 · Multimodal 64 · Reasoning 66

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoningthinkingdefaultSourceTrend
GPQA Diamond90.8%#33Epoch AI Benchmarking Hub
LMArena Hard Prompts1485#48LMArena
LiveBench Reasoningnot in index79.4%#40LiveBench
GPQA Diamond (AA)not in index78.8%#20191.1%#50Artificial Analysis
Humanity's Last Exam (AA)not in index19.6%#16337.5%#69Artificial Analysis
GPQA Diamond (Vals)not in index89.1%#30Vals AI

Coding

Benchmarkno reasoningthinkingdefaultSourceTrend
SWE-bench Verified (Epoch)76.7%#9Epoch AI Benchmarking Hub
SciCode53.5%#41SciCode
WeirdML55.9%#55WeirdML
LMArena Coding1514#36LMArena
LMArena WebDev1509#38LMArena
LiveBench Codingnot in index78.6%#20LiveBench
SciCode (AA)not in index51.5%#59Artificial Analysis
LiveCodeBench86.8%#19Vals AI
SWE-bench (Vals)not in index76.2%#39Vals AI

Agents & tools

Benchmarkno reasoningthinkingdefaultSourceTrend
OSWorld-Verified 2.04.6%#12OSWorld
APEX-Agents18.9%#38Mercor
LiveBench Agentic Codingnot in index46.9%#36LiveBench
Terminal-Bench Hard37.9%#6143.9%#35Artificial Analysis
τ²-Bench Telecom (AA)not in index93.9%#3995.9%#16Artificial Analysis
Terminal-Bench 2.1 (Vals)53.6%#43Vals AI
HiL-Bench18.7%#14Scale AI SEAL

Maths

Benchmarkno reasoningthinkingdefaultSourceTrend
FrontierMath Tiers 1–357.2%#31Epoch AI Benchmarking Hub
FrontierMath Tier 425.6%#35Epoch AI Benchmarking Hub
OTIS Mock AIME96.1%#25Epoch AI Benchmarking Hub
ProofBench16.0%#41Vals AI
LiveBench Mathematicsnot in index84.3%#43LiveBench
AIME 202695.8%#13MathArena
HMMT February 202694.7%#8MathArena
MathArena Apex24.0%#13MathArena

Knowledge

Benchmarkno reasoningthinkingdefaultSourceTrend
SimpleQA Verified34.9%#42Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index65.1%#45LiveBench
AA-Omniscience-9.2#1345.3#71Artificial Analysis
MMLU-Pro87.6%#25Vals AI
LegalBench84.7%#27Vals AI
CorpFin66.7%#16Vals AI
TaxEval74.7%#32Vals AI

Instruction following

Benchmarkno reasoningthinkingdefaultSourceTrend
LiveBench Languagenot in index75.1%#37LiveBench
IFBench44.3%#21076.0%#23Artificial Analysis

Human preference

Benchmarkno reasoningthinkingdefaultSourceTrend
LMArena Text1461#49LMArena
EQ-Bench 41202#16EQ-Bench

Multimodal

Benchmarkno reasoningthinkingdefaultSourceTrend
LMArena Vision1281#27LMArena
MMMU-Pro79.4%#42Artificial Analysis

Long context

Benchmarkno reasoningthinkingdefaultSourceTrend
AA-LCR69.7%#16881.0%#37Artificial Analysis

Composite

Benchmarkno reasoningthinkingdefaultSourceTrend
Epoch Capabilities Indexnot in index151.0#40Epoch AI Benchmarking Hub
LiveBench70.5%#41LiveBench
AA Intelligence Index23.6#12931.3#72Artificial Analysis
Vals Indexnot in index43.5#32Vals AI

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
CoreWeave219 tok/s0.33 s$0.650$3.41262kfp4
Inceptron80 tok/s0.56 s$0.590$2.45262kint4
Parasail79 tok/s1.09 s$0.750$3.50262kint4
Decart71 tok/s0.49 s$0.587$2.47262kfp4
Crusoe66 tok/s0.44 s$0.700$3.50262kbf16
DigitalOcean43 tok/s1.14 s$0.950$4.00262k
AtlasCloud40 tok/s1.77 s$0.950$4.00262kint4
Baidu Qianfan38 tok/s1.49 s$0.580$2.44262kfp4
StreamLake37 tok/s1.70 s$0.599$2.52256kfp8
NovitaAI36 tok/s1.79 s$0.800$3.40262k
Moonshot AI33 tok/s2.88 s$0.950$4.00262kint4
Cloudflare32 tok/s0.71 s$0.950$4.00262k
Chutes31 tok/s2.75 s$0.580$3.40262kint4
Venice25 tok/s1.30 s$0.750$3.50256kint4
SiliconFlow23 tok/s2.03 s$0.770$3.40262kfp8
Phala21 tok/s2.11 s$1.09$4.60262k
DeepInfra20 tok/s1.20 s$0.750$3.50262kfp4
GMICloud6 tok/s6.11 s$0.855$3.60262kfp8

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$1.10$2.20$3.30$4.40May 26Jun 26Jul 26Aug 26
input outputnow $0.580 in · $2.44 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.160 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0016$0.00131.6 min
Summarise a 30-page report12,000 / 600$0.014$0.00671.7 min
Code edit6,000 / 1,500$0.012$0.00812.0 min
Agentic coding session60,000 / 4,000$0.073$0.0372.8 min
Structured extraction2,000 / 200$0.0027$0.00151.5 min

See also

Data as of 9 Sept 2026. Compare these configurations.