BenchLeader

DeepSeek V4 Flash

Best configuration ranks #106 of 610 on the BenchLeader Index at 59.3 ±3.1 (high reasoning effort). Last measured 8 Sept 2026. Released 31 Jul 2026.

Blended price
$0.168/M
$0.130 in · $0.280 out
Output speed
118 tok/s
First answer
18 s
first token 1.26 s
Context
1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index59
  2. Reasoning61
  3. Coding61
  4. Agents & tools60
  5. Knowledge54
  6. Instruction following71
  7. Long context62
  8. Composite60

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning48.4#325118 tok/s18 s$0.0001Agents & tools 63 · Composite 53 · Instruction following 49 · Knowledge 43 · Long context 46 · Reasoning 33
low118 tok/s18 s$0.0001Reasoning 57
highbest59.3#106118 tok/s18 s$0.0001Agents & tools 60 · Coding 61 · Composite 60 · Instruction following 71 · Knowledge 54 · Long context 62 · Reasoning 61
max54.7#193118 tok/s18 s$0.0001Coding 58 · Knowledge 45 · Maths 57 · Reasoning 64
default59.1#109118 tok/s18 s$0.0001Agents & tools 60 · Coding 48 · Composite 61 · Human preference 63 · Instruction following 76 · Knowledge 57 · Long context 66 · Maths 61 · Reasoning 58

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoninglowhighmaxdefaultSourceTrend
GPQA Diamond91.0%#27Epoch AI Benchmarking Hub
SimpleBench61.1%#26SimpleBench
LMArena Hard Prompts1459#81LMArena
LiveBench Reasoningnot in index86.6%#23LiveBench
GPQA Diamond (AA)not in index71.6%#27586.7%#10790.8%#55Artificial Analysis
Humanity's Last Exam (AA)not in index7.8%#28630.3%#10138.5%#63Artificial Analysis
GPQA Diamond (Vals)not in index89.9%#27Vals AI
Kagi LLM Benchmark52.2%#76Kagi LLM Benchmark
ARC-AGI-111.8%#16784.0%#7787.0%#6389.0%#56ARC Prize
ARC-AGI-22.1%#14146.0%#7556.0%#6461.4%#53ARC Prize

Coding

Benchmarkno reasoninglowhighmaxdefaultSourceTrend
SciCode49.9%#62SciCode
WeirdML57.0%#5363.0%#40WeirdML
FrontierCode18.8%#24Cognition
LMArena Coding1483#86LMArena
LiveBench Codingnot in index75.0%#38LiveBench
SciCode (AA)not in index40.2%#11450.4%#67Artificial Analysis
LiveCodeBench87.3%#16Vals AI
SWE-bench (Vals)not in index88.8%#11Vals AI

Agents & tools

Benchmarkno reasoninglowhighmaxdefaultSourceTrend
LMArena Agent2.2#19LMArena
LiveBench Agentic Codingnot in index46.8%#37LiveBench
Terminal-Bench Hard34.1%#8738.6%#5935.6%#72Artificial Analysis
τ²-Bench Telecom (AA)not in index94.4%#3195.6%#2095.0%#27Artificial Analysis
Terminal-Bench 2.1 (Vals)67.0%#28Vals AI

Maths

Benchmarkno reasoninglowhighmaxdefaultSourceTrend
FrontierMath Tiers 1–357.5%#30Epoch AI Benchmarking Hub
FrontierMath Tier 424.4%#36Epoch AI Benchmarking Hub
OTIS Mock AIME94.4%#39Epoch AI Benchmarking Hub
ProofBench56.0%#13Vals AI
LiveBench Mathematicsnot in index86.8%#36LiveBench
AIME 202695.8%#13MathArena
HMMT February 202693.9%#10MathArena
MathArena Apex27.1%#11MathArena

Knowledge

Benchmarkno reasoninglowhighmaxdefaultSourceTrend
SimpleQA Verified33.6%#48Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index79.3%#8LiveBench
AA-Omniscience-44.0#278-23.1#193-14.3#159Artificial Analysis
MMLU-Pro86.2%#45Vals AI
LegalBench77.7%#99Vals AI
CorpFin61.9%#52Vals AI
TaxEval70.7%#85Vals AI

Instruction following

Benchmarkno reasoninglowhighmaxdefaultSourceTrend
LiveBench Languagenot in index79.2%#28LiveBench
IFBench47.2%#18773.5%#4079.2%#11Artificial Analysis

Human preference

Benchmarkno reasoninglowhighmaxdefaultSourceTrend
LMArena Text1436#86LMArena

Long context

Benchmarkno reasoninglowhighmaxdefaultSourceTrend
AA-LCR41.7%#28272.0%#14079.7%#59Artificial Analysis

Composite

Benchmarkno reasoninglowhighmaxdefaultSourceTrend
Epoch Capabilities Indexnot in index154.5#28Epoch AI Benchmarking Hub
LiveBench74.2%#32LiveBench
AA Intelligence Index18.9#18624.8#11834.5#55Artificial Analysis
Vals Indexnot in index53.6#22Vals AI

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Phala73 tok/s1.44 s$0.200$0.4001.0M
NextBit72 tok/s1.46 s$0.150$0.3501.0Mfp8
Baidu Qianfan68 tok/s0.77 s$0.068$0.1361.0Mfp8
SiliconFlow68 tok/s1.69 s$0.130$0.2801.0Mfp8
Alibaba Cloud Int.64 tok/s0.92 s$0.134$0.2681Mfp8
GMICloud56 tok/s2.06 s$0.091$0.1821.0Mfp8
NovitaAI45 tok/s1.38 s$0.140$0.2801.0Mfp8
Azure (US)44 tok/s1.23 s$0.210$0.5601.0M
Mancer40 tok/s0.80 s$0.190$0.5001.0Mfp8
AtlasCloud39 tok/s1.09 s$0.140$0.2801.0Mfp4
Parasail37 tok/s0.87 s$0.140$0.2801.0Mfp8
Venice34 tok/s1.26 s$0.138$0.2751M
StreamLake22 tok/s8.82 s$0.070$0.1401.0Mfp8
DeepInfra16 tok/s2.50 s$0.090$0.1801.0Mfp8
DigitalOcean11 tok/s1.12 s$0.068$0.1681.0M

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$0.085$0.169$0.254$0.339May 26Jun 26Jul 26Aug 26
input outputnow $0.086 in · $0.171 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.003 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0001$0.000120.7 s
Summarise a 30-page report12,000 / 600$0.0017$0.000623.3 s
Code edit6,000 / 1,500$0.0012$0.000630.9 s
Agentic coding session60,000 / 4,000$0.0089$0.003252.2 s
Structured extraction2,000 / 200$0.0003$0.000119.9 s

See also

Data as of 9 Sept 2026. Compare these configurations.