Gemini 3.8 Live scores 97.7% on Big Bench Audio, second only to Qwen
ArtificialAnlys · x · 2026-09-16
Audio reasoning data from Artificial Analysis: Gemini 3.8 Live Extended Thinking (High) scores 97.7% on Big Bench Audio, ahead of Grok Voice Think Fast 2.0 High (97.2%) and behind only Qwen Audio 3.0 Realtime Plus (99.2%). The standard model scores 91.7%.
Related event: Gemini 3.8 Live Tops Voice Benchmarks at Lowest Price(6 posts)→
More from Models
- Stealth Startup Unveils Non-Chat Model Claiming 100x Speed and Cost Edge Over LLMs — hardimanjames · 2026-09-16
- Analyst argues Meta could actually exclude user data from training, unlike OpenAI's lawyerly wording — ivan_bezdomny · 2026-09-16
- DeepSeek V4.1 Flash Keeps Timing Out on 2-Hour Agentic Benchmarks, Author Shares Failure Logs — sebnadeau · 2026-09-16
- GPT-6 Astra remakes Game of Thrones in low-poly Blender after 8h autonomous run — OWazabi · 2026-09-16
- Why Meta Could Actually Keep Your Agent Data Out of Training — and Frontier Labs Can't — ivan_bezdomny · 2026-09-16
- TypeSafe AI launches Jev, a model for fast structured decisions with confidence scores — yogthinks · 2026-09-16