BenchLeader

DeepSeek V4 Pro

Best configuration ranks #95 of 610 on the BenchLeader Index at 60.0 ±3.0 (high reasoning effort). Last measured 8 Sept 2026. Released 24 Apr 2026.

Blended price
$0.548/M
$0.440 in · $0.870 out
Output speed
66 tok/s
First answer
32 s
first token 1.39 s
Context
1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index60
  2. Reasoning62
  3. Agents & tools64
  4. Knowledge59
  5. Instruction following69
  6. Long context62
  7. Composite67

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning48.8#31670 tok/s1.75 s$0.0004Agents & tools 65 · Coding 32 · Composite 55 · Instruction following 47 · Knowledge 50 · Long context 52 · Maths 45 · Reasoning 41
low79 tok/s27 s$0.0004Reasoning 61
highbest60.0#9566 tok/s32 s$0.0004Agents & tools 64 · Composite 67 · Instruction following 69 · Knowledge 59 · Long context 62 · Reasoning 62
max56.1#16379 tok/s27 s$0.0004Agents & tools 44 · Coding 61 · Knowledge 59 · Maths 57 · Reasoning 64
default59.7#9879 tok/s27 s$0.0004Agents & tools 53 · Coding 56 · Composite 67 · Human preference 59 · Instruction following 74 · Knowledge 64 · Long context 67 · Maths 58 · Reasoning 58

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoninglowhighmaxdefaultSourceTrend
GPQA Diamond73.2%#14591.7%#24Epoch AI Benchmarking Hub
LMArena Hard Prompts1480#53LMArena
LiveBench Reasoningnot in index85.8%#25LiveBench
GPQA Diamond (AA)not in index71.7%#27490.5%#5992.8%#28Artificial Analysis
Humanity's Last Exam (AA)not in index8.3%#28035.2%#7941.0%#50Artificial Analysis
GPQA Diamond (Vals)not in index92.4%#16Vals AI
Kagi LLM Benchmark53.5%#70Kagi LLM Benchmark
ARC-AGI-113.0%#16590.5%#5087.2%#6190.0%#53ARC Prize
ARC-AGI-20.8%#16456.3%#6359.7%#5961.3%#54ARC Prize

Coding

Benchmarkno reasoninglowhighmaxdefaultSourceTrend
SWE-bench Verified (Epoch)77.6%#6Epoch AI Benchmarking Hub
SciCode50.0%#60SciCode
WeirdML66.2%#36WeirdML
FrontierCode17.6%#25Cognition
LMArena Coding1502#61LMArena
LMArena WebDev1446#55LMArena
LiveBench Codingnot in index77.2%#28LiveBench
SciCode (AA)not in index51.0%#62Artificial Analysis
LiveCodeBench87.5%#13Vals AI
IOI35.8%#15Vals AI
SWE-bench (Vals)not in index96.4%#2Vals AI

Agents & tools

Benchmarkno reasoninglowhighmaxdefaultSourceTrend
LMArena Agent4.4#12-0.4#26LMArena
LiveBench Agentic Codingnot in index55.0%#22LiveBench
Terminal-Bench Hard36.4%#6941.7%#4846.2%#30Artificial Analysis
τ²-Bench Telecom (AA)not in index91.2%#5994.2%#3496.2%#15Artificial Analysis
Terminal-Bench 2.1 (Vals)54.7%#4050.2%#48Vals AI

Maths

Benchmarkno reasoninglowhighmaxdefaultSourceTrend
FrontierMath Tiers 1–364.6%#24Epoch AI Benchmarking Hub
FrontierMath Tier 426.8%#31Epoch AI Benchmarking Hub
OTIS Mock AIME46.7%#16998.6%#13Epoch AI Benchmarking Hub
ProofBench16.0%#4150.0%#17Vals AI
LiveBench Mathematicsnot in index95.1%#8LiveBench
AIME 202696.7%#6MathArena
HMMT February 202693.9%#10MathArena
MathArena Apex28.1%#10MathArena

Knowledge

Benchmarkno reasoninglowhighmaxdefaultSourceTrend
SimpleQA Verified52.9%#17Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index79.2%#11LiveBench
AA-Omniscience-29.9#214-10.6#1440.8#89Artificial Analysis
MMLU-Pro87.3%#32Vals AI
LegalBench82.4%#65Vals AI
CorpFin65.4%#28Vals AI
TaxEval73.1%#53Vals AI

Instruction following

Benchmarkno reasoninglowhighmaxdefaultSourceTrend
LiveBench Languagenot in index82.1%#19LiveBench
IFBench45.8%#19771.3%#5776.5%#20Artificial Analysis

Human preference

Benchmarkno reasoninglowhighmaxdefaultSourceTrend
LMArena Text1458#54LMArena
EQ-Bench 41166#17EQ-Bench

Long context

Benchmarkno reasoninglowhighmaxdefaultSourceTrend
AA-LCR53.0%#24370.3%#15780.3%#45Artificial Analysis

Composite

Benchmarkno reasoninglowhighmaxdefaultSourceTrend
Epoch Capabilities Indexnot in index155.4#21Epoch AI Benchmarking Hub
LiveBench77.4%#15LiveBench
AA Intelligence Index20.8#16530.1#8136.3#50Artificial Analysis
Vals Indexnot in index52.4#25Vals AI

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
NovitaAI67 tok/s1.39 s$1.60$3.201.0Mfp8
Baseten62 tok/s0.60 s$1.74$3.481.0Mfp4
Parasail62 tok/s0.87 s$1.74$3.481.0Mfp8
Baseten (US)62 tok/s0.58 s$1.74$3.481.0Mfp4
Baidu Qianfan54 tok/s1.15 s$0.869$1.741.0Mfp8
Azure (US)47 tok/s1.40 s$1.91$3.831.0M
NextBit46 tok/s2.15 s$1.80$3.551.0Mfp8
GMICloud41 tok/s2.13 s$0.957$1.911.0Mfp8
StreamLake39 tok/s1.54 s$0.870$1.741.0Mfp8
SiliconFlow37 tok/s2.15 s$1.50$3.131.0Mfp8
AtlasCloud37 tok/s1.31 s$1.68$3.381.0Mfp4
DigitalOcean35 tok/s0.80 s$0.870$1.741.0M
Alibaba Cloud Int.33 tok/s1.66 s$1.42$2.831Mfp8
Venice28 tok/s1.75 s$1.65$3.301M
DeepInfra27 tok/s1.31 s$1.30$2.601.0Mfp8

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$0.957$1.91$2.87$3.83May 26Jun 26Jul 26Aug 26
input outputnow $0.870 in · $1.74 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.004 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0004$0.000336.6 s
Summarise a 30-page report12,000 / 600$0.0058$0.001941.2 s
Code edit6,000 / 1,500$0.0039$0.002054.9 s
Agentic coding session60,000 / 4,000$0.030$0.0101.5 min
Structured extraction2,000 / 200$0.0011$0.000435.1 s

See also

Data as of 9 Sept 2026. Compare these configurations.