Gemini 3.1 Flash Live tops new Speech Agent Arena; GPT-Realtime-1.5 leads task success
_philschmid · x · 2026-08-22
Artificial Analysis launched its Speech Agent Arena leaderboard, ranking speech-to-speech models by human preference in blind live voice conversations like booking dental appointments and ordering takeout. Google's Gemini 3.1 Flash Live Preview (Minimal) ranks #1 with Elo 1046 and a 74.6% task success rate; the High variant is second (Elo 1014).
OpenAI's GPT-Realtime-1.5 posts the highest task success rate at 85.1% (Elo 1000, #3), with ElevenLabs Agents at 90.5%. Amazon Nova 2.0 Sonic, Grok Voice, and Qwen3.5 Omni also make the board—Qwen3.5 Omni Flash manages only a 29.1% task success rate.
Related event: Speech Agent Arena launches with Gemini leading human preference rankings(3 posts)→
More from Models
- NVIDIA's AVO harness lifts Opus 5 from 30% to 100% on ARC-AGI-3 — daniel_mac8 · 2026-08-22
- Qwen 3.8 vs 3.6: Low reasoning mode loops less — Lair98 · 2026-08-22
- GLM-5.3 Takes 2nd Place on Creative Writing Benchmark with Qualitative Analysis — zero0_one1 · 2026-08-22
- Private Benchmark: 0x Alpha Underperforms on Low Reasoning Tasks — toptickcrypto · 2026-08-22
- Ox Alpha Model Test: Surprising Performance in Code and Data Analysis — killerbee1432 · 2026-08-22
- Qwen3.8-27B: 3x Faster Long-Context Decoding with DFlash2 + XQA — stargate425 · 2026-08-22