Gemini 3.1 Flash Live tops new Speech Agent Arena; GPT-Realtime-1.5 leads task success

_philschmid · x · 2026-08-22

Artificial Analysis launched its Speech Agent Arena leaderboard, ranking speech-to-speech models by human preference in blind live voice conversations like booking dental appointments and ordering takeout. Google's Gemini 3.1 Flash Live Preview (Minimal) ranks #1 with Elo 1046 and a 74.6% task success rate; the High variant is second (Elo 1014).

OpenAI's GPT-Realtime-1.5 posts the highest task success rate at 85.1% (Elo 1000, #3), with ElevenLabs Agents at 90.5%. Amazon Nova 2.0 Sonic, Grok Voice, and Qwen3.5 Omni also make the board—Qwen3.5 Omni Flash manages only a 29.1% task success rate.

Related event: Speech Agent Arena launches with Gemini leading human preference rankings(3 posts)→

Original post →

More from Models

Models channel →