Google's Gemini 3.8 Live Speech-to-Speech Models Top Voice Rankings at $0.84/Hour Input

DeepLearningAI · x · 2026-10-07

DeepLearning.AI breaks down Google's new Gemini 3.8 Live speech-to-speech models: single-system listening, reasoning, and responding without relay-style handoffs. The Extended Thinking version ranks first on Artificial Analysis' Speech-to-Speech Index; the standard version ranks second in human-judged blind conversations and costs $0.84 per hour of input audio, the lowest in the index. Both accept image and video input.

Original post →

More from Models

Models channel →