BenchLeader

Best models for high-volume extraction

Millions of short, near-identical calls: pull the fields out of a document, sort a ticket, tag a review. Nobody is waiting on any single response, so latency barely matters and frontier quality is wasted money. Price and throughput decide it, with quality as a floor — the job still has to come out right.

As of 11 Oct 2026, DeepSeek V4 Flash at max effort fits this best, with price per 1m of $0.168 and output speed of 222.

How this is weighted

  • 45%Price per 1Mat this volume the bill is the decision; a tenfold price difference dwarfs every other consideration
  • 25%Output speedthroughput sets how long the queue takes to drain
  • 20%Instruction followingstructured output that holds its shape every time is worth more here than cleverness
  • 10%Qualitysome quality signal, but only enough to keep obviously weak models out

A floor, not a ranking: quality must reach 40 — cheap and fast is worthless if the extraction is wrong.

203 of 434 ranked models qualify: the rest either miss one of the measurements the weighting depends on, or fall outside a limit above. We would rather leave a model out than score it on the dimensions it happens to have.

#ModelFitPrice per 1MOutput speedInstruction followingQualityContext window
1DeepSeek V4 FlashmaxDeepSeek93$0.1682227760.61M
2Ling 3.0 Flash FinAnt Group92$0.115321–54.4262k
3Ling 3.0 FlashAnt Group91$0.115334–52.9262k
4NewClaude Haiku 5.5maxAnthropic89$0.200240–57.91M
5Nemotron 3.5 LightningNVIDIA88$0.095307–47.61M
6Ling 3.0 Flash VLAnt Group88$0.115145–55.7262k
7Granite 4.2 3BIBM87$0.053215–45.0131k
8Nemotron 3 Nano 30B A3BthinkingNVIDIA85$0.0882197046.6131k
9GPT-6 LunamaxOpenAI85$0.200139–58.11.1M
10Step 3.5 FlashStepFun83$0.1501606654.0256k
11Gemma 4 12BthinkingGoogle82$0.1501147250.5262k
12DeepSeek V4.1 FlashmaxDeepSeek81$0.525217–61.81M
13gpt-oss-20bhighOpenAI80$0.0981756544.1131k
14GPT-5 nanohighOpenAI79$0.1381246747.6400k
15GPT-5.4 NanoxhighOpenAI79$0.4631647453.5400k
16Mercury 2.5Inception79$0.375818–45.5260k
17gpt-oss-120bhighOpenAI79$0.2621846848.6131k
18Mercury 2Inception78$0.3754566947.4128k
19Celeris-1Celeris78$0.3251523–41.5131k
20Deepseek v4 Flash VisionmaxDeepSeek77$0.660221–58.41M
21Solar Pro 3Upstage77$0.2621587043.2131k
22Qwen3 30B A3BAlibaba77$0.215119–48.541k
23Gemini 2.5 Flash-LitethinkingGoogle76$0.1752895146.11.0M
24Nemotron 3 Super 120B A12bthinkingNVIDIA76$0.4501527050.7262k
25GPT-5.6 LunamaxOpenAI76$0.450113–58.71.1M
26Inkling SmallThinking Machines76$0.525160–55.1524k
27MiniMax M3MiniMax75$0.525928056.51M
28Gemini 3 FlashthinkingGoogle75$1.132087659.21.0M
29Command A+Cohere75$0.6001757250.6192k
30Step 3.7 FlashStepFun74$0.4381526651.6256k
31Gemini 3.1 Flash LitehighGoogle74$0.563323–48.91.0M
32GPT-5 minihighOpenAI74$0.6881297354.2400k
33Gemini 3.5 Flash LiteGoogle74$0.850365–54.71.0M
34Trinity LargethinkingArcee AI74$0.3883395747.0512k
35Qwen3.5 FlashAlibaba74$0.17579–52.21M
36MiMo-V2.6-FlashXiaomi73$0.17558–59.71.0M
37Nemotron 3 Nano Omni 30B A3bthinkingNVIDIA73$0.4502406345.1256k
38Qwen3.5-35B-A3BthinkingAlibaba73$0.6881437152.8262k
39Nemotron 3 Ultra 550B A55BthinkingNVIDIA72$1.071427956.61M
40MiniMax-M2MiniMax71$0.525897154.5205k

How to read this

Fit is a percentile blend, not a score out of a hundred: each dimension is ranked against every other model that could be judged here, then combined with the weights above. It says how well a model matches this job compared with the alternatives — a model can fit voice work superbly and sit well down the quality leaderboard, which is the point of ranking by job rather than by index.

Disagree with the weighting? That is a reasonable thing to do, which is why it is printed rather than hidden. To set your own constraints instead, use the model finder.

Cite as: BenchLeader, “Best models for high-volume extraction”, https://www.benchleader.com/use/bulk-extraction, data as of 11 Oct 2026.