BenchLeader

VoiceCodeBench

Speech recognition of entities across eight workflow domains. Run by Vals AI.

As of 19 Sept 2026, GPT Live Transcribe leads VoiceCodeBench on BenchLeader with 67.7%, ahead of Ink 2 at 62.0%, across 18 model configurations with a published result.

Published by
Vals AI
Category
Multimodal
Index weight
Reference only
Models
18
Data as of
19 Sept 2026

Vals AI (vals.ai).

What the test looks like

Speech-recognition models transcribe audio containing domain entities such as drug names or ticker symbols; scored on getting those entities right.

How it is scored

Entity accuracy, run by Vals AI.

What to keep in mind

A speech benchmark: only speech models are listed, and it does not measure reasoning.

18 of 18
#
1GPT Live TranscribeOpenAI67.7%2026-09-01
2Ink 2Unknown62.0%2026-09-01
3GPT 4o TranscribeOpenAI59.7%2026-09-01
4Nova 3 StreamingAmazon59.3%2026-09-01
5Inworld Stt 1Unknown57.3%2026-09-01
6Muse Voice TranscribeMeta57.0%2026-09-01
7Whisper Large v3 TurboUnknown57.0%2026-09-01
8Flux General EnUnknown55.7%2026-09-01
9Chirp 3Unknown55.3%2026-09-01
10Scribe v2 RealtimeUnknown55.3%2026-09-01
11Whisper Large v3Unknown54.3%2026-09-01
12Voxtral Mini 2602Mistral AI53.0%2026-09-01
13Grok SttxAI52.0%2026-09-01
14GPT 4o Mini Transcribe Streaming ResponseOpenAI51.7%2026-09-01
15Voxtral Mini Transcribe Realtime 2602Mistral AI46.0%2026-09-01
16Transcribe 03 2026Cohere44.7%2026-09-01
17Universal Language ModelUnknown37.3%2026-09-01
18Universal 3.5 ProUnknown30.3%2026-09-01

Cite as: BenchLeader, “VoiceCodeBench leaderboard”, https://www.benchleader.com/benchmarks/vals_voice_code_bench, data as of 19 Sept 2026.

VoiceCodeBench: questions

What does VoiceCodeBench measure?
Speech-recognition models transcribe audio containing domain entities such as drug names or ticker symbols; scored on getting those entities right. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads VoiceCodeBench?
GPT Live Transcribe leads VoiceCodeBench with 67.7% as of 19 Sept 2026, ahead of Ink 2 at 62.0%.
How many models have VoiceCodeBench results?
18 model configurations have a VoiceCodeBench result on BenchLeader, all taken from Vals AI.
Who runs VoiceCodeBench and how often is it updated?
VoiceCodeBench is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
Does VoiceCodeBench count toward the BenchLeader Index?
No. VoiceCodeBench is shown for reference but left out of the composite index.