BenchLeader

DeepSeek V4 Flash 0731

DeepSeek V4 Flash 0731 is a DeepSeek open-weights reasoning model, released 31 Jul 2026. Its best configuration (high reasoning effort) ranks #164 of 372 on the BenchLeader Index at 56.3 ±3.1, in the upper half. It scores highest in coding (61) and lowest in knowledge (55). At $0.094 per million tokens blended it is among the cheapest fifth of ranked models. Output speed of 37 tokens per second puts it in the slowest quarter, with a first token in 0.8 s. It has been measured at 5 reasoning-effort settings; tables show the best-scoring one. Last measured 10 Sept 2026.

Blended price
$0.094/M
$0.065 in · $0.180 out
Output speed
37 tok/s
OpenRouter traffic, 7-day median; not yet measured by Artificial Analysis
First answer
0.84 s
Context
1.3M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index56
  2. Reasoning61
  3. Coding61
  4. Knowledge55

Versions

DeepSeek has shipped 2 models under this name. Each is ranked on its own results; a newer version often has fewer results so far, which holds its index nearer the average until more arrive.

ModelReleasedIndexRank
DeepSeek V4 Flash 0731highthis page31 Jul 202656.3#164
DeepSeek V4 Flashhigh24 Apr 202659.1#110

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning37 tok/s0.84 s$0.0001Reasoning 33
low37 tok/s0.84 s$0.0001Reasoning 57
highbest56.3#16437 tok/s0.84 s$0.0001Coding 61 · Knowledge 55 · Reasoning 61
max54.5#20137 tok/s0.84 s$0.0001Coding 59 · Knowledge 44 · Maths 57 · Reasoning 64
default51.1#27737 tok/s0.84 s$0.0001Agents & tools 55 · Coding 33 · Composite 49 · Maths 61 · Reasoning 60

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoninglowhighmaxdefaultSource
GPQA Diamond91.0%#28Epoch AI Benchmarking Hub
SimpleBench61.1%#26SimpleBench
LiveBench Reasoningnot in index86.6%#24LiveBench
GPQA Diamond (Vals)not in index89.9%#27Vals AI
ARC-AGI-111.8%#16784.0%#7787.0%#6389.0%#56ARC Prize
ARC-AGI-22.1%#14146.0%#7556.0%#6461.4%#53ARC Prize

Coding

Benchmarkno reasoninglowhighmaxdefaultSource
SciCode49.9%#62SciCode
WeirdML57.0%#5363.0%#40WeirdML
FrontierCode18.8%#24Cognition
LiveBench Codingnot in index75.0%#39LiveBench
LiveCodeBench87.3%#17Vals AI
SWE-bench (Vals)not in index88.8%#11Vals AI

Agents & tools

Benchmarkno reasoninglowhighmaxdefaultSource
LiveBench Agentic Codingnot in index46.8%#38LiveBench
Terminal-Bench 2.1 (Vals)67.0%#29Vals AI

Maths

Benchmarkno reasoninglowhighmaxdefaultSource
FrontierMath Tiers 1–357.5%#31Epoch AI Benchmarking Hub
FrontierMath Tier 424.4%#37Epoch AI Benchmarking Hub
OTIS Mock AIME94.4%#41Epoch AI Benchmarking Hub
ProofBench56.0%#13Vals AI
LiveBench Mathematicsnot in index86.8%#38LiveBench

Knowledge

Benchmarkno reasoninglowhighmaxdefaultSource
SimpleQA Verified33.6%#50Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index79.3%#8LiveBench
MMLU-Pro86.2%#46Vals AI
LegalBench77.7%#100Vals AI
CorpFin61.9%#52Vals AI
TaxEval70.7%#85Vals AI

Instruction following

Benchmarkno reasoninglowhighmaxdefaultSource
LiveBench Languagenot in index79.2%#29LiveBench

Composite

Benchmarkno reasoninglowhighmaxdefaultSource
Epoch Capabilities Indexnot in index154.5#28Epoch AI Benchmarking Hub
LiveBench74.2%#33LiveBench
Vals Indexnot in index53.6#23Vals AI

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Baidu Qianfan143 tok/s0.68 s$0.440$1.321.0Mfp8
Reka AI141 tok/s0.59 s$0.110$0.660262kfp4
Wafer124 tok/s0.45 s$0.100$0.2501.0M
NovitaAI81 tok/s1.60 s$0.409$1.231.0Mfp8
Fireworks79 tok/s0.82 s$0.220$0.6601.0M
Alibaba Cloud Int.65 tok/s1.04 s$0.176$0.5281M
Makora64 tok/s0.48 s$0.090$0.1951M
CoreWeave64 tok/s0.44 s$0.130$0.280262kfp8
Inceptron60 tok/s0.53 s$0.130$0.3501.0Mfp4
Together53 tok/s0.68 s$0.140$0.2801.0M
GMICloud47 tok/s2.24 s$0.286$0.8581.0Mfp4
AtlasCloud43 tok/s1.90 s$0.440$1.321.0Mfp4
SiliconFlow42 tok/s1.65 s$0.220$0.6601.0Mfp8
StreamLake40 tok/s2.04 s$0.088$0.2641.0Mfp8
Parasail39 tok/s0.61 s$0.140$0.2801.0Mfp8
Cloudflare37 tok/s1.48 s$0.440$1.321.3M
Baseten (US)36 tok/s1.77 s$0.130$0.2601.0Mfp8
NextBit36 tok/s2.13 s$0.352$1.061.0Mfp8
Sail Research30 tok/s1.46 s$0.065$0.1801.0Mfp4
Baseten29 tok/s1.55 s$0.130$0.2601.0Mfp8
Phala28 tok/s0.92 s$0.440$1.321.0M
Mancer26 tok/s0.74 s$0.165$0.5001.0Mfp8
Venice25 tok/s0.98 s$0.175$0.3501M
Relace23 tok/s1.41 s$0.065$0.1801.0Mfp4
DigitalOcean23 tok/s0.55 s$0.080$0.2521.0M
OpenInference21 tok/s1.64 s$0.050$0.1601.0Mfp8
DeepInfra21 tok/s0.84 s$0.060$0.1801.0Mfp8
Morph4 tok/s1.94 s$0.123$0.3471.0Mbf16

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$0.204$0.408$0.611$0.815Aug 26Sept 26
input outputnow $0.088 in · $0.264 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.040 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0001$0.00019.0 s
Summarise a 30-page report12,000 / 600$0.0009$0.000717.1 s
Code edit6,000 / 1,500$0.0007$0.000541.4 s
Agentic coding session60,000 / 4,000$0.0046$0.00351.8 min
Structured extraction2,000 / 200$0.0002$0.00016.2 s

See also

Data as of 10 Sept 2026. Compare these configurations.