Gemini 3.8 Live tops Tau Voice agentic benchmark at 68.6%, up from 37.7%
ArtificialAnlys · x · 2026-09-16
Agentic data from Artificial Analysis: Gemini 3.8 Live Extended Thinking (High) takes #1 on the Tau Voice benchmark at 68.6%, ahead of GPT-Live-1 (Astra) at 67.9% and Grok at 56.5% — up from just 37.7% for the previous generation. The standard model also ranks #2 in Speech Agent Arena preference (Elo 1083) with 93.2% task success.
Related event: Gemini 3.8 Live Tops Voice Benchmarks at Lowest Price(6 posts)→
More from Models
- Stealth Startup Unveils Non-Chat Model Claiming 100x Speed and Cost Edge Over LLMs — hardimanjames · 2026-09-16
- TypeSafe launches Jev: a structured-probability model that judges agent outputs in 0.7 seconds — hardimanjames · 2026-09-16
- Analyst argues Meta could actually exclude user data from training, unlike OpenAI's lawyerly wording — ivan_bezdomny · 2026-09-16
- DeepSeek V4.1 Flash Keeps Timing Out on 2-Hour Agentic Benchmarks, Author Shares Failure Logs — sebnadeau · 2026-09-16
- GPT-6 Astra remakes Game of Thrones in low-poly Blender after 8h autonomous run — OWazabi · 2026-09-16
- Why Meta Could Actually Keep Your Agent Data Out of Training — and Frontier Labs Can't — ivan_bezdomny · 2026-09-16