Grok Voice tops Speech Agent Arena with 94.6% task success rate, beating Gemini and GPT
XFreeze · x · 2026-09-17
Grok Voice "Think Fast 2.0" from xAI ranks #1 on Artificial Analysis' Speech Agent Arena with a 94.6% task success rate, ahead of Gemini 3.8 Live, GPT-Realtime-2.1 High, and other tested voice agents.
The benchmark uses real people talking to hidden voice agents on practical tasks, then verifies whether the AI understood the request, used the right tools, and actually completed the job — measuring task completion rather than just human-likeness.
More from Models
- Jev passes 8/9 computer-use tasks, makes decisions 13.6x faster than Astra/Codex — iamrobotbear · 2026-09-17
- Typesafe AI Launches Jev: A Decision Engine at 70-500ms and $0.042 Per Million Tokens — brandon_galang · 2026-09-17
- GPT-6 Astra tops Terminal-Bench 4.0 at 57.7%, Claude Opus 5 hits 51.8% at half the price — shensi · 2026-09-17
- Dev take: models are smart enough now — focus on making them fail less — rickasaurus · 2026-09-17
- OpenAI Publishes Misalignment Disclosure Framework, Plus Six Incident Reports From Six Months of Training — Thom_Wolf · 2026-09-17
- A 'Quadrillion-Parameter' Model Surfaces on X, But Details Remain Unverified — KyeGomezB · 2026-09-17