BenchLeader

Best models for voice agents

In a spoken conversation the thing people feel is the pause before the model starts talking. A model that scores two points higher but thinks for four seconds first is the worse choice, every time. So latency leads here, tool use comes next because voice agents almost always have to do something, and quality is a bar to clear rather than a thing to rank on.

As of 11 Oct 2026, Gemini 2.5 Flash-Lite fits this best, with time to first answer of 0.56 s and output speed of 168.

How this is weighted

  • 45%Time to first answerthe silence before the first word is what a caller notices; this figure waits through any reasoning phase, so a model that thinks before answering is penalised exactly as a listener would penalise it
  • 20%Output speedspeech synthesis reads the stream as it arrives, so throughput sets whether the voice keeps up or stutters mid-sentence
  • 15%Tool usea voice agent that cannot call a tool is a voice that cannot book, look up or cancel anything
  • 10%Instruction followingstaying in character, keeping answers short and following a script matter more here than raw intelligence
  • 10%Whole-turn timea turn that never ends is its own failure mode, however fast it started

A hard limit: time to first answer over 2s is excluded entirely — a pause past about a second already reads as a dropped call, and past two seconds the conversation has broken down — a model that slow is not a slow option for voice, it is not an option.

A floor, not a ranking: quality must reach 45 — below roughly average quality the conversation stops being useful however quick it is.

65 of 434 ranked models qualify: the rest either miss one of the measurements the weighting depends on, or fall outside a limit above. We would rather leave a model out than score it on the dimensions it happens to have.

#ModelFitTime to first answerOutput speedTool useInstruction followingPrice per 1M
1Gemini 2.5 Flash-LiteGoogle810.56 s1684952$0.175
2Gemini 3.5 FlashminimalGoogle780.97 s1917449$3.38
3Grok 4.20no reasoningSpaceXAI780.59 s935351$3.00
4Nova 2 LiteAmazon770.58 s13040–$0.935
5Gemini 2.5 Flashno reasoningGoogle770.46 s1914442$0.850
6GPT-4.1 miniOpenAI770.73 s1254655$0.700
7Gemini 3 Flashno reasoningGoogle751.04 s2066156$1.13
8Claude Haiku 4.5no reasoningAnthropic750.47 s855745$2.00
9GPT-4.1OpenAI750.86 s1335845$3.50
10GLM 5.2Zhipu AI740.89 s9555–$2.15
11DeepSeek V4.1 FlashhighDeepSeek721.01 s9861–$0.262
12Qwen3 32BAlibaba720.39 s2759–$1.23
13GPT-5.6 Terrano reasoningOpenAI710.77 s9144–$4.50
14Nemotron 3 Ultra 550B A55BNVIDIA710.42 s7040–$1.00
15Grok 4.3no reasoningSpaceXAI690.78 s1024149$1.56
16Mistral Medium 3Mistral AI690.66 s1294343$0.800
17Qwen3.6 35B A3BAlibaba680.64 s8535–$0.557
18GPT-5.5no reasoningOpenAI670.90 s875448$11.25
19Gemma 4 26B A4BthinkingGoogle670.56 s484271$0.168
20GPT-5.4no reasoningOpenAI670.89 s716750$5.63
21Gemini 3.7 FlashlowGoogle671.29 s29258–$1.50
22Qwen3.8 27BAlibaba660.69 s5645–$1.06
23gpt-oss-120bOpenAI660.39 s762347$0.262
24Qwen3.5-9Bno reasoningAlibaba640.70 s814941$0.190
25GPT-5.4 MinihighOpenAI621.01 s9944–$1.69
26DeepSeek V4 Flash 0731DeepSeek610.79 s4255–$0.175
27Claude Sonnet 5no reasoningAnthropic600.77 s6144–$4.00
28GPT-5.1no reasoningOpenAI591.03 s885346$3.44
29GPT-5.2no reasoningOpenAI580.82 s734049$4.81
30GPT-5.4 NanohighOpenAI570.97 s8732–$0.463
31Kimi K2.6Moonshot AI550.85 s5037–$1.71
32GPT-5.6 Solno reasoningOpenAI521.09 s7848–$8.00
33GPT-5minimalOpenAI521.12 s804948$3.44
34Gemini 3.6 FlashGoogle521.30 s12347–$1.50
35Claude Sonnet 4Anthropic501.19 s485456$6.00
36Claude Sonnet 4.6lowAnthropic500.97 s427145$6.00
37MiniMax-M1MiniMax490.94 s593644$0.963
38Claude Opus 4.5no reasoningAnthropic481.14 s476945$10.00
39Claude Opus 4.8highAnthropic471.48 s6263–$10.00
40DeepSeek V4 FlashhighDeepSeek461.53 s555972$0.168

Announced, not yet measured

This model is too new for any of our publishers to have run it. The figures below are the maker’s own, so they cannot be ranked against measured ones — but anyone picking a model for this job should know it exists.

  • Mercury VoiceInception
    Time to first answer 0.32 s

    No measured output speed, tool use, instruction following, whole-turn time yet, so no fit score.

How to read this

Fit is a percentile blend, not a score out of a hundred: each dimension is ranked against every other model that could be judged here, then combined with the weights above. It says how well a model matches this job compared with the alternatives — a model can fit voice work superbly and sit well down the quality leaderboard, which is the point of ranking by job rather than by index.

Disagree with the weighting? That is a reasonable thing to do, which is why it is printed rather than hidden. To set your own constraints instead, use the model finder.

Cite as: BenchLeader, “Best models for voice agents”, https://www.benchleader.com/use/voice, data as of 11 Oct 2026.