Gemini 3.8 Live debuts #2 in Speech Agent Arena with Elo 1083, 93.2% task success
ArtificialAnlys · x · 2026-09-16
Speech Agent Arena data from Artificial Analysis: in blind live voice preference tests, Gemini 3.8 Live debuts #2 at Elo 1083, behind only Gemini 3.1 Flash Live (1096) and ahead of GPT-Live-1 (Sol, 1053). It completes 93.2% of tasks, second to Grok's 94.6%. The Extended Thinking variant sits at Elo 990 with 89.1% task success.
Related event: Gemini 3.8 Live Tops Voice Benchmarks at Lowest Price(6 posts)→
More from Models
- That 1M-token context window can really burn your bill — peterjliu · 2026-09-16
- Speculation: top open-weight models may be distilling OpenAI and Anthropic, missing training code hints — dan_s_becker · 2026-09-16
- Science LLM benchmarks have flawed answers; fixing them significantly raises model scores — Profanion · 2026-09-16
- Aaronson hears AI companies have cracked longstanding TCS open problems, sitting on major announcements — scaling01 · 2026-09-16
- Prior Labs releases TabPFN-3.5, a tabular foundation model handling 1M rows, tops benchmarks — tuanacelik · 2026-09-16
- User reports Sonnet 5 fails at shop price comparisons and hallucinates local stores — RachelVT42 · 2026-09-16