Best models for high-volume extraction
Millions of short, near-identical calls: pull the fields out of a document, sort a ticket, tag a review. Nobody is waiting on any single response, so latency barely matters and frontier quality is wasted money. Price and throughput decide it, with quality as a floor — the job still has to come out right.
As of 11 Oct 2026, DeepSeek V4 Flash at max effort fits this best, with price per 1m of $0.168 and output speed of 222.
How this is weighted
- 45%Price per 1Mat this volume the bill is the decision; a tenfold price difference dwarfs every other consideration
- 25%Output speedthroughput sets how long the queue takes to drain
- 20%Instruction followingstructured output that holds its shape every time is worth more here than cleverness
- 10%Qualitysome quality signal, but only enough to keep obviously weak models out
A floor, not a ranking: quality must reach 40 — cheap and fast is worthless if the extraction is wrong.
203 of 434 ranked models qualify: the rest either miss one of the measurements the weighting depends on, or fall outside a limit above. We would rather leave a model out than score it on the dimensions it happens to have.
| # | Model | Fit | Price per 1M | Output speed | Instruction following | Quality | Context window |
|---|---|---|---|---|---|---|---|
| 1 | 93 | $0.168 | 222 | 77 | 60.6 | 1M | |
| 2 | 92 | $0.115 | 321 | – | 54.4 | 262k | |
| 3 | 91 | $0.115 | 334 | – | 52.9 | 262k | |
| 4 | New | 89 | $0.200 | 240 | – | 57.9 | 1M |
| 5 | 88 | $0.095 | 307 | – | 47.6 | 1M | |
| 6 | 88 | $0.115 | 145 | – | 55.7 | 262k | |
| 7 | 87 | $0.053 | 215 | – | 45.0 | 131k | |
| 8 | 85 | $0.088 | 219 | 70 | 46.6 | 131k | |
| 9 | 85 | $0.200 | 139 | – | 58.1 | 1.1M | |
| 10 | 83 | $0.150 | 160 | 66 | 54.0 | 256k | |
| 11 | 82 | $0.150 | 114 | 72 | 50.5 | 262k | |
| 12 | 81 | $0.525 | 217 | – | 61.8 | 1M | |
| 13 | 80 | $0.098 | 175 | 65 | 44.1 | 131k | |
| 14 | 79 | $0.138 | 124 | 67 | 47.6 | 400k | |
| 15 | 79 | $0.463 | 164 | 74 | 53.5 | 400k | |
| 16 | 79 | $0.375 | 818 | – | 45.5 | 260k | |
| 17 | 79 | $0.262 | 184 | 68 | 48.6 | 131k | |
| 18 | 78 | $0.375 | 456 | 69 | 47.4 | 128k | |
| 19 | Celeris-1Celeris | 78 | $0.325 | 1523 | – | 41.5 | 131k |
| 20 | 77 | $0.660 | 221 | – | 58.4 | 1M | |
| 21 | 77 | $0.262 | 158 | 70 | 43.2 | 131k | |
| 22 | 77 | $0.215 | 119 | – | 48.5 | 41k | |
| 23 | 76 | $0.175 | 289 | 51 | 46.1 | 1.0M | |
| 24 | 76 | $0.450 | 152 | 70 | 50.7 | 262k | |
| 25 | 76 | $0.450 | 113 | – | 58.7 | 1.1M | |
| 26 | Inkling SmallThinking Machines | 76 | $0.525 | 160 | – | 55.1 | 524k |
| 27 | 75 | $0.525 | 92 | 80 | 56.5 | 1M | |
| 28 | 75 | $1.13 | 208 | 76 | 59.2 | 1.0M | |
| 29 | 75 | $0.600 | 175 | 72 | 50.6 | 192k | |
| 30 | 74 | $0.438 | 152 | 66 | 51.6 | 256k | |
| 31 | 74 | $0.563 | 323 | – | 48.9 | 1.0M | |
| 32 | 74 | $0.688 | 129 | 73 | 54.2 | 400k | |
| 33 | 74 | $0.850 | 365 | – | 54.7 | 1.0M | |
| 34 | 74 | $0.388 | 339 | 57 | 47.0 | 512k | |
| 35 | 74 | $0.175 | 79 | – | 52.2 | 1M | |
| 36 | 73 | $0.175 | 58 | – | 59.7 | 1.0M | |
| 37 | 73 | $0.450 | 240 | 63 | 45.1 | 256k | |
| 38 | 73 | $0.688 | 143 | 71 | 52.8 | 262k | |
| 39 | 72 | $1.07 | 142 | 79 | 56.6 | 1M | |
| 40 | 71 | $0.525 | 89 | 71 | 54.5 | 205k |
How to read this
Fit is a percentile blend, not a score out of a hundred: each dimension is ranked against every other model that could be judged here, then combined with the weights above. It says how well a model matches this job compared with the alternatives — a model can fit voice work superbly and sit well down the quality leaderboard, which is the point of ranking by job rather than by index.
Disagree with the weighting? That is a reasonable thing to do, which is why it is printed rather than hidden. To set your own constraints instead, use the model finder.
Cite as: BenchLeader, “Best models for high-volume extraction”, https://www.benchleader.com/use/bulk-extraction, data as of 11 Oct 2026.