BenchLeader

DeepSeek V4 Pro 0813

DeepSeek V4 Pro 0813 is a DeepSeek open-weights reasoning model, released 13 Aug 2026. Its best configuration (max reasoning effort) ranks #160 of 372 on the BenchLeader Index at 56.4 ±3.4, in the upper half. It scores highest in reasoning (64) and lowest in agents & tools (44). At $0.869 per million tokens blended it is mid-priced. Output speed of 20 tokens per second puts it in the slowest quarter, with a first token in 1.4 s. It has been measured at 5 reasoning-effort settings; tables show the best-scoring one. Last measured 10 Sept 2026.

Blended price
$0.869/M
$0.580 in · $1.74 out
Output speed
20 tok/s
OpenRouter traffic, 7-day median; not yet measured by Artificial Analysis
First answer
1.43 s
Context
1.0M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index56
  2. Reasoning64
  3. Coding61
  4. Agents & tools44
  5. Maths60
  6. Knowledge59

Versions

DeepSeek has shipped 2 models under this name. Each is ranked on its own results; a newer version often has fewer results so far, which holds its index nearer the average until more arrive.

ModelReleasedIndexRank
DeepSeek V4 Pro 0813maxthis page13 Aug 202656.4#160
DeepSeek V4 Prohigh24 Apr 202659.5#102

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning20 tok/s1.43 s$0.0008Reasoning 32
low20 tok/s1.43 s$0.0008Reasoning 61
high20 tok/s1.43 s$0.0008Reasoning 62
maxbest56.4#16020 tok/s1.43 s$0.0008Agents & tools 44 · Coding 61 · Knowledge 59 · Maths 60 · Reasoning 64
default20 tok/s1.43 s$0.0017Composite 59 · Maths 58

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoninglowhighmaxdefaultSource
GPQA Diamond91.7%#25Epoch AI Benchmarking Hub
LiveBench Reasoningnot in index85.8%#26LiveBench
GPQA Diamond (Vals)not in index92.4%#16Vals AI
ARC-AGI-113.0%#16590.5%#5087.2%#6190.0%#53ARC Prize
ARC-AGI-20.8%#16456.3%#6359.7%#5961.3%#54ARC Prize

Coding

Benchmarkno reasoninglowhighmaxdefaultSource
SciCode49.2%#65SciCode
WeirdML66.2%#36WeirdML
LiveBench Codingnot in index77.2%#29LiveBench
LiveCodeBench87.5%#13Vals AI
SWE-bench (Vals)not in index96.4%#2Vals AI

Agents & tools

Benchmarkno reasoninglowhighmaxdefaultSource
LiveBench Agentic Codingnot in index55.0%#23LiveBench
Terminal-Bench 2.1 (Vals)54.7%#41Vals AI

Maths

Benchmarkno reasoninglowhighmaxdefaultSource
FrontierMath Tiers 1–364.6%#25Epoch AI Benchmarking Hub
FrontierMath Tier 426.8%#32Epoch AI Benchmarking Hub
OTIS Mock AIME98.6%#14Epoch AI Benchmarking Hub
ProofBench50.0%#17Vals AI
LiveBench Mathematicsnot in index95.1%#8LiveBench

Knowledge

Benchmarkno reasoninglowhighmaxdefaultSource
SimpleQA Verified52.9%#17Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index79.2%#12LiveBench
MMLU-Pro87.0%#36Vals AI
LegalBench82.4%#66Vals AI
CorpFin65.4%#28Vals AI
TaxEval73.1%#52Vals AI

Instruction following

Benchmarkno reasoninglowhighmaxdefaultSource
LiveBench Languagenot in index82.1%#19LiveBench

Composite

Benchmarkno reasoninglowhighmaxdefaultSource
Epoch Capabilities Indexnot in index155.4#21Epoch AI Benchmarking Hub
LiveBench77.4%#16LiveBench
Vals Indexnot in index52.4#26Vals AI

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
CoreWeave124 tok/s0.76 s$1.31$3.961.0Mfp8
Together110 tok/s0.53 s$1.32$3.961.0M
Ionstream97 tok/s1.06 s$1.10$3.301.0M
Parasail72 tok/s0.59 s$1.32$3.961.0Mfp8
NovitaAI68 tok/s1.44 s$0.990$2.971.0Mfp8
DeepInfra63 tok/s1.06 s$1.30$2.601.0Mfp8
Cloudflare60 tok/s1.26 s$1.32$3.961.0M
StreamLake59 tok/s3.32 s$0.660$1.981.0M
Alibaba Cloud Int.57 tok/s1.44 s$0.581$1.741M
GMICloud56 tok/s2.31 s$1.06$3.171.0Mfp8
Phala55 tok/s1.61 s$1.45$4.361.0M
Baseten (US)53 tok/s0.67 s$1.32$3.961.0Mfp4
Baidu Qianfan52 tok/s1.14 s$0.580$1.741.0Mfp8
SiliconFlow49 tok/s1.50 s$1.32$3.961.0Mfp8
Fireworks47 tok/s1.86 s$1.32$3.961.0M
Sail Research44 tok/s1.04 s$1.32$3.961.0Mfp4
Baseten37 tok/s0.50 s$1.32$3.961.0Mfp4
NextBit34 tok/s4.04 s$1.12$3.371.0Mfp8
DigitalOcean30 tok/s0.79 s$1.32$3.961.0M
DeepSeek20 tok/s1.43 s$0.660$1.981.0M

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$1.09$2.18$3.27$4.36Aug 26
input outputnow $0.660 in · $1.98 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.100 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0008$0.000616.4 s
Summarise a 30-page report12,000 / 600$0.0080$0.003731.4 s
Code edit6,000 / 1,500$0.0061$0.00391.3 min
Agentic coding session60,000 / 4,000$0.042$0.0203.4 min
Structured extraction2,000 / 200$0.0015$0.000811.4 s

See also

Data as of 10 Sept 2026. Compare these configurations.