Best models for voice agents
In a spoken conversation the thing people feel is the pause before the model starts talking. A model that scores two points higher but thinks for four seconds first is the worse choice, every time. So latency leads here, tool use comes next because voice agents almost always have to do something, and quality is a bar to clear rather than a thing to rank on.
As of 11 Oct 2026, Gemini 2.5 Flash-Lite fits this best, with time to first answer of 0.56 s and output speed of 168.
How this is weighted
- 45%Time to first answerthe silence before the first word is what a caller notices; this figure waits through any reasoning phase, so a model that thinks before answering is penalised exactly as a listener would penalise it
- 20%Output speedspeech synthesis reads the stream as it arrives, so throughput sets whether the voice keeps up or stutters mid-sentence
- 15%Tool usea voice agent that cannot call a tool is a voice that cannot book, look up or cancel anything
- 10%Instruction followingstaying in character, keeping answers short and following a script matter more here than raw intelligence
- 10%Whole-turn timea turn that never ends is its own failure mode, however fast it started
A hard limit: time to first answer over 2s is excluded entirely — a pause past about a second already reads as a dropped call, and past two seconds the conversation has broken down — a model that slow is not a slow option for voice, it is not an option.
A floor, not a ranking: quality must reach 45 — below roughly average quality the conversation stops being useful however quick it is.
65 of 434 ranked models qualify: the rest either miss one of the measurements the weighting depends on, or fall outside a limit above. We would rather leave a model out than score it on the dimensions it happens to have.
| # | Model | Fit | Time to first answer | Output speed | Tool use | Instruction following | Price per 1M |
|---|---|---|---|---|---|---|---|
| 1 | 81 | 0.56 s | 168 | 49 | 52 | $0.175 | |
| 2 | 78 | 0.97 s | 191 | 74 | 49 | $3.38 | |
| 3 | 78 | 0.59 s | 93 | 53 | 51 | $3.00 | |
| 4 | 77 | 0.58 s | 130 | 40 | – | $0.935 | |
| 5 | 77 | 0.46 s | 191 | 44 | 42 | $0.850 | |
| 6 | 77 | 0.73 s | 125 | 46 | 55 | $0.700 | |
| 7 | 75 | 1.04 s | 206 | 61 | 56 | $1.13 | |
| 8 | 75 | 0.47 s | 85 | 57 | 45 | $2.00 | |
| 9 | 75 | 0.86 s | 133 | 58 | 45 | $3.50 | |
| 10 | 74 | 0.89 s | 95 | 55 | – | $2.15 | |
| 11 | 72 | 1.01 s | 98 | 61 | – | $0.262 | |
| 12 | 72 | 0.39 s | 27 | 59 | – | $1.23 | |
| 13 | 71 | 0.77 s | 91 | 44 | – | $4.50 | |
| 14 | 71 | 0.42 s | 70 | 40 | – | $1.00 | |
| 15 | 69 | 0.78 s | 102 | 41 | 49 | $1.56 | |
| 16 | 69 | 0.66 s | 129 | 43 | 43 | $0.800 | |
| 17 | 68 | 0.64 s | 85 | 35 | – | $0.557 | |
| 18 | 67 | 0.90 s | 87 | 54 | 48 | $11.25 | |
| 19 | 67 | 0.56 s | 48 | 42 | 71 | $0.168 | |
| 20 | 67 | 0.89 s | 71 | 67 | 50 | $5.63 | |
| 21 | 67 | 1.29 s | 292 | 58 | – | $1.50 | |
| 22 | 66 | 0.69 s | 56 | 45 | – | $1.06 | |
| 23 | 66 | 0.39 s | 76 | 23 | 47 | $0.262 | |
| 24 | 64 | 0.70 s | 81 | 49 | 41 | $0.190 | |
| 25 | 62 | 1.01 s | 99 | 44 | – | $1.69 | |
| 26 | 61 | 0.79 s | 42 | 55 | – | $0.175 | |
| 27 | 60 | 0.77 s | 61 | 44 | – | $4.00 | |
| 28 | 59 | 1.03 s | 88 | 53 | 46 | $3.44 | |
| 29 | 58 | 0.82 s | 73 | 40 | 49 | $4.81 | |
| 30 | 57 | 0.97 s | 87 | 32 | – | $0.463 | |
| 31 | 55 | 0.85 s | 50 | 37 | – | $1.71 | |
| 32 | 52 | 1.09 s | 78 | 48 | – | $8.00 | |
| 33 | 52 | 1.12 s | 80 | 49 | 48 | $3.44 | |
| 34 | 52 | 1.30 s | 123 | 47 | – | $1.50 | |
| 35 | 50 | 1.19 s | 48 | 54 | 56 | $6.00 | |
| 36 | 50 | 0.97 s | 42 | 71 | 45 | $6.00 | |
| 37 | 49 | 0.94 s | 59 | 36 | 44 | $0.963 | |
| 38 | 48 | 1.14 s | 47 | 69 | 45 | $10.00 | |
| 39 | 47 | 1.48 s | 62 | 63 | – | $10.00 | |
| 40 | 46 | 1.53 s | 55 | 59 | 72 | $0.168 |
Announced, not yet measured
This model is too new for any of our publishers to have run it. The figures below are the maker’s own, so they cannot be ranked against measured ones — but anyone picking a model for this job should know it exists.
Mercury VoiceInception
Time to first answer 0.32 sNo measured output speed, tool use, instruction following, whole-turn time yet, so no fit score.
How to read this
Fit is a percentile blend, not a score out of a hundred: each dimension is ranked against every other model that could be judged here, then combined with the weights above. It says how well a model matches this job compared with the alternatives — a model can fit voice work superbly and sit well down the quality leaderboard, which is the point of ranking by job rather than by index.
Disagree with the weighting? That is a reasonable thing to do, which is why it is printed rather than hidden. To set your own constraints instead, use the model finder.
Cite as: BenchLeader, “Best models for voice agents”, https://www.benchleader.com/use/voice, data as of 11 Oct 2026.