BenchLeader

AudioMC

Multiple-choice questions asked and answered in audio. Scale AI.

As of 19 Sept 2026, Gemini 3.8 Flash leads AudioMC on BenchLeader with 60.4%, ahead of Inkling Small at 54.9%, across 27 model configurations with a published result.

Published by
Scale AI SEAL
Category
Multimodal
Index weight
Reference only
Models
27
Data as of
19 Sept 2026

Scale AI SEAL Leaderboards.

What the test looks like

Questions are spoken, sometimes with audio evidence, and the model must answer from listening.

How it is scored

Accuracy, published by Scale AI.

What to keep in mind

Only audio-capable models are listed.

27 of 27
#
1Gemini 3.8 FlashhighGoogle60.4%64.52026-09-09
2Inkling SmallThinking Machinesopen ↗54.9%56.32026-07-30
3Gemini 3 ProthinkingGoogle54.6%2025-12-17
4GPT Realtime 2xhighOpenAI48.5%2026-05-07
5Gemini 2.5 ProthinkingGoogle46.9%2025-12-17
6Tml Interaction SmallUnknown43.4%2026-05-11
7Gemini 2.5 FlashthinkingGoogle40.0%48.42025-12-17
8GPT Realtime 2OpenAI37.6%2026-05-07
9Gemini 3.1 Flash LivethinkingGoogle36.1%2026-03-26
10GPT Realtime 1.5OpenAI34.7%2026-02-23
11Gemini 3.1 Flash LiveGoogle26.8%2026-03-26
12Voxtral Small 24B 2507Mistral AI26.3%2026-02-25
13Gemini 2.5 FlashGoogle26.1%52.52026-02-23
14GPT 4o AudioOpenAI25.4%2026-02-25
15Qwen3 Omni 30B A3B InstructAlibabaopen ↗24.3%39.42026-02-25
16GPT RealtimeOpenAI23.4%2025-12-17
17Gemini 2.5 Flash Native Audio 12 2025thinkingGoogle21.5%2026-02-25
18Mimo Audio 7BthinkingUnknown19.7%2025-12-17
19Mimo Audio 7BUnknown18.6%2025-12-17
20GPT Realtime MiniOpenAI16.6%2026-02-23
21Gemma 3n E4BGoogleopen15.5%36.42025-12-17
22Phi 4 MultimodalMicrosoftopen15.5%2025-12-17
23GPT 4o Mini AudioOpenAI14.8%2025-12-17
24Gemini 2.5 Flash Native Audio 12 2025no reasoningGoogle13.9%2026-03-14
25Kimi Audio 7BMoonshot AI13.7%2025-12-17
26Qwen2.5-Omni 7BAlibabaopen ↗11.9%2025-12-17
27Lfm2 Audio 1 5BLiquid AI9.3%2025-12-17

Cite as: BenchLeader, “AudioMC leaderboard”, https://www.benchleader.com/benchmarks/scale_audiomc, data as of 19 Sept 2026.

AudioMC: questions

What does AudioMC measure?
Questions are spoken, sometimes with audio evidence, and the model must answer from listening. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads AudioMC?
Gemini 3.8 Flash leads AudioMC with 60.4% as of 19 Sept 2026, ahead of Inkling Small at 54.9%.
How many models have AudioMC results?
27 model configurations have a AudioMC result on BenchLeader, all taken from Scale AI SEAL.
Who runs AudioMC and how often is it updated?
AudioMC is published by Scale AI SEAL. BenchLeader re-reads the published results every morning and records the date each result was published.
Does AudioMC count toward the BenchLeader Index?
No. AudioMC is shown for reference but left out of the composite index.