VoiceCodeBench
Speech recognition of entities across eight workflow domains. Run by Vals AI.
As of 19 Sept 2026, GPT Live Transcribe leads VoiceCodeBench on BenchLeader with 67.7%, ahead of Ink 2 at 62.0%, across 18 model configurations with a published result.
- Published by
- Vals AI
- Category
- Multimodal
- Index weight
- Reference only
- Models
- 18
- Data as of
- 19 Sept 2026
Vals AI (vals.ai).
What the test looks like
Speech-recognition models transcribe audio containing domain entities such as drug names or ticker symbols; scored on getting those entities right.
How it is scored
Entity accuracy, run by Vals AI.
What to keep in mind
A speech benchmark: only speech models are listed, and it does not measure reasoning.
- 1GPT Live Transcribe67.7%
- 2Ink 262.0%
- 3GPT 4o Transcribe59.7%
- 4Nova 3 Streaming59.3%
- 5Inworld Stt 157.3%
- 6Muse Voice Transcribe57.0%
- 7Whisper Large v3 Turbo57.0%
- 8Flux General En55.7%
- 9Chirp 355.3%
- 10Scribe v2 Realtime55.3%
- 11Whisper Large v354.3%
- 12Voxtral Mini 260253.0%
- 13Grok Stt52.0%
- 14GPT 4o Mini Transcribe Streaming Response51.7%
- 15Voxtral Mini Transcribe Realtime 260246.0%
18 of 18
| # | ||||
|---|---|---|---|---|
| 1 | 67.7% | – | 2026-09-01 | |
| 2 | Ink 2Unknown | 62.0% | – | 2026-09-01 |
| 3 | 59.7% | – | 2026-09-01 | |
| 4 | 59.3% | – | 2026-09-01 | |
| 5 | Inworld Stt 1Unknown | 57.3% | – | 2026-09-01 |
| 6 | 57.0% | – | 2026-09-01 | |
| 7 | Whisper Large v3 TurboUnknown | 57.0% | – | 2026-09-01 |
| 8 | Flux General EnUnknown | 55.7% | – | 2026-09-01 |
| 9 | Chirp 3Unknown | 55.3% | – | 2026-09-01 |
| 10 | Scribe v2 RealtimeUnknown | 55.3% | – | 2026-09-01 |
| 11 | Whisper Large v3Unknown | 54.3% | – | 2026-09-01 |
| 12 | 53.0% | – | 2026-09-01 | |
| 13 | 52.0% | – | 2026-09-01 | |
| 14 | 51.7% | – | 2026-09-01 | |
| 15 | 46.0% | – | 2026-09-01 | |
| 16 | 44.7% | – | 2026-09-01 | |
| 17 | Universal Language ModelUnknown | 37.3% | – | 2026-09-01 |
| 18 | Universal 3.5 ProUnknown | 30.3% | – | 2026-09-01 |
Cite as: BenchLeader, “VoiceCodeBench leaderboard”, https://www.benchleader.com/benchmarks/vals_voice_code_bench, data as of 19 Sept 2026.
VoiceCodeBench: questions
- What does VoiceCodeBench measure?
- Speech-recognition models transcribe audio containing domain entities such as drug names or ticker symbols; scored on getting those entities right. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads VoiceCodeBench?
- GPT Live Transcribe leads VoiceCodeBench with 67.7% as of 19 Sept 2026, ahead of Ink 2 at 62.0%.
- How many models have VoiceCodeBench results?
- 18 model configurations have a VoiceCodeBench result on BenchLeader, all taken from Vals AI.
- Who runs VoiceCodeBench and how often is it updated?
- VoiceCodeBench is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does VoiceCodeBench count toward the BenchLeader Index?
- No. VoiceCodeBench is shown for reference but left out of the composite index.