Speech Agent Arena Launch: Gemini Leads Preference, Grok Leads Success Rate
ArtificialAnlys · x · 2026-08-21
ArtificialAnlys announced the Speech Agent Arena, evaluating Speech-to-Speech models via human interactions in real-world scenarios (e.g., ordering takeout). It measures both Conversational Preference (Elo) and Task Success Rate.
Key Results:
- Preference: Google Gemini 3.1 Flash Live Preview - Minimal leads at 1,046 Elo, followed by GPT-Realtime-1.5 at 1,000. Preferred models tend to respond faster (lower TTFA) and sound more natural.
- Task Success: Grok Voice Think Fast 2.0 High leads at 94.7%, followed by GPT-Realtime-2.1 High at 91.5%. Notably, Gemini's top preference score (1,046 Elo) contrasts with its 74.6% success rate, indicating that pleasant conversation does not always result in successful task completion.
More from Models
- V4-Flash-Vision tested: 2x faster, cheaper and better than V4-Flash-0731 — cedric_chee · 2026-08-21
- Kimi K3.1 quietly testing on Code Arena; Ox Alpha said to be Zhipu's GLM 5.3 Flash — i_dg23 · 2026-08-21
- OpenRouter's stealth model Ox Alpha sparks guessing game over its identity — AccBalanced · 2026-08-21
- ThursdAI weekly: Qwen 27B and GLM 5.3 beat GPTs; OpenAI pauses RL for security — thursdai_pod · 2026-08-21
- Ox Alpha's stunning fluid sim 'one-shot' appears to be a copy of an existing GitHub repo — scaling01 · 2026-08-21
- Creator's GPT Image 3 wishlist: perfect pixel art, zero noise, working spritesheets — Angaisb_ · 2026-08-21