Gemini 3.8 Live tops Tau Voice agentic benchmark at 68.6%, up from 37.7%

ArtificialAnlys · x · 2026-09-16

Agentic data from Artificial Analysis: Gemini 3.8 Live Extended Thinking (High) takes #1 on the Tau Voice benchmark at 68.6%, ahead of GPT-Live-1 (Astra) at 67.9% and Grok at 56.5% — up from just 37.7% for the previous generation. The standard model also ranks #2 in Speech Agent Arena preference (Elo 1083) with 93.2% task success.

Related event: Gemini 3.8 Live Tops Voice Benchmarks at Lowest Price(6 posts)→

Original post →

More from Models

Models channel →