New Leaderboard Reveals the State of Voice AI, Highlighting Multilingual Gaps
The rapid advancement of voice AI is outpacing the tools used to evaluate it. As major AI labs – OpenAI, Google DeepMind, Anthropic, and xAI – race to deploy models capable of natural, real-time conversation, a new benchmark is emerging to assess their performance through genuine human interaction. Scale AI, the data annotation startup, has launched Voice Showdown, a global, preference-based arena designed to benchmark voice AI and provide free access to leading frontier models.
Addressing the Limitations of Existing Benchmarks
Current voice AI evaluations often rely on synthetic speech, English-only prompts, and scripted test sets that don’t reflect real-world usage. Scale AI’s Voice Showdown aims to overcome these limitations by utilizing spontaneous voice conversations across over 60 languages. “Voice AI is really the fastest moving frontier in AI right now,” said Janie Gu, product manager for Showdown at Scale AI. “But the way that we evaluate voice models hasn’t kept up.”
How Voice Showdown Works
Built on Scale’s ChatLab platform, Voice Showdown allows users to interact with high-tier AI models—typically requiring multiple paid subscriptions—at no cost. In exchange, users participate in blind, head-to-head “battles,” choosing which of two anonymized voice models provides a better experience. This data contributes to a human-preference leaderboard.
The platform currently offers two evaluation modes:
- Dictate: Users speak, and models respond with text.
- Speech-to-Speech (S2S): Users speak, and models respond with speech.
A third mode, Full Duplex – capturing real-time, interruptible conversation – is under development.
Incentivized Voting and Controlled Comparisons
Voice Showdown differentiates itself from other benchmarks, like LM Arena, through its incentive structure. After voting for a preferred model, users are switched to that model for the remainder of their conversation, encouraging more thoughtful responses. The system also controls for potential biases by ensuring simultaneous response streaming, matching voice gender, and anonymizing model identities during voting.
Current Leaderboard Results (March 18, 2026)
As of March 18, 2026, the leaderboard includes 11 frontier models evaluated across 52 model-voice pairs.
Dictate Leaderboard (Speech-In, Text-Out)
- Gemini 3 Pro (1073)
- Gemini 3 Flash (1068)
- GPT-4o Audio (1019)
- Question 3 Omni (1000)
- Voxtral Little (925)
- Gemma 3n (918)
- GPT Realtime (875)
- Phi-4 Multimodal (729)
Gemini 3 Pro and Gemini 3 Flash are statistically tied for the top rank.
Speech-to-Speech (S2S) Leaderboard
- Gemini 2.5 Flash Audio (1060)
- GPT-4o Audio (1059)
- Grok Voice (1024)
- Question 3 Omni (1000)
- GPT Realtime (962)
- GPT Realtime 1.5 (920)
Gemini 2.5 Flash Audio and GPT-4o Audio are statistically tied for the top rank in baseline evaluations. After adjusting for style controls, GPT-4o Audio leads with an Elo score of 1,102.
Key Findings: Multilingual Performance and Voice Selection
The data reveals significant gaps in multilingual capabilities. GPT Realtime 1.5 responds in English to non-English prompts roughly 20% of the time, even for widely supported languages like Hindi, Spanish, and Turkish. Gemini 2.5 Flash Audio and GPT-4o Audio exhibit a lower mismatch rate of approximately 7%.
the study highlights the importance of voice selection within a single model. The best-performing voice from one unnamed model won 30 percentage points more often than its worst-performing voice, despite sharing the same underlying reasoning and generation backend.
The Impact of Conversational Length
Voice Showdown also assessed model performance over extended conversations. Content quality becomes a more significant failure point as conversations progress, accounting for 43% of failures after 11 turns, compared to 23% on the first turn. GPT Realtime variants showed marginal improvement with longer contexts.
Looking Ahead
Scale AI plans to introduce Full Duplex evaluation to capture the dynamics of real-time, interruptible conversations. The Voice Showdown leaderboard is available at scale.com/showdown, with a public waitlist open for access to ChatLab.