Gemini 3.8 Live scores 97.7% on Big Bench Audio, second only to Qwen

ArtificialAnlys · x · 2026-09-16

Audio reasoning data from Artificial Analysis: Gemini 3.8 Live Extended Thinking (High) scores 97.7% on Big Bench Audio, ahead of Grok Voice Think Fast 2.0 High (97.2%) and behind only Qwen Audio 3.0 Realtime Plus (99.2%). The standard model scores 91.7%.

Related event: Gemini 3.8 Live Tops Voice Benchmarks at Lowest Price(6 posts)→

Original post →

More from Models

Models channel →