FULL STORY

Gemini 3.8 Live Launches, Tops Voice AI Benchmarks

Google launched Gemini 3.8 Live and its Extended Thinking variant with real-time voice support in 97 languages. Subsequent Artificial Analysis benchmarks showed the model topping multiple voice categories at the lowest price.

2026-09-15 ~ 2026-09-16 · 2 episodes · 27 posts

Episode 1 · Google Launches Gemini 3.8 Live and Extended Thinking, Topping Speech Quality Index (2026-09-15, 21 posts)

Google officially launched two real-time audio conversation models on September 16—Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking—calling them its most advanced Gemini audio models to date, achieving SOTA performance at frontier-model pricing and performance. Developers can try them online in Google AI Studio.

Confirmed

  • Both models support 97 languages with seamless mid-conversation switching, automatic language detection, and barge-in during voice conversations
  • Shared features include upgraded reasoning, near-real-time visual understanding, and background (asynchronous) tool calls that don't interrupt the chat
  • Gemini 3.8 Live Extended Thinking adds extended thinking; DeepMind's official demo showed it acting as a coding tutor, reasoning and explaining without breaking the conversation
  • According to philschmid, Gemini 3.8 Live tops the Artificial Analysis Quality Index with a score of 82.6
  • The day before launch (September 15), testingcatalog cited Bedros Pamboukian's finding that entries for both models had appeared in the GCP console quotas and metrics pages, corroborating the pre-launch leak

Why it matters

  • Continuing the real-time voice agent line started by last month's Gemini 3.5 Transcribe, it is seen as a direct answer to GPT Live (per testingcatalog)
  • The combined ability to "speak, think, and run background tasks without interrupting the user's flow" marks the evolution of voice assistants from pure conversation toward real-time agents that can execute tasks

1 more related posts →

Episode 2 · Gemini 3.8 Live Tops Voice Benchmarks at Lowest Price (2026-09-16, 6 posts)

On 09-16, Artificial Analysis published evaluation results for Gemini 3.8 Live across multiple voice capabilities, showing the model leading or near the top in latency, agent capability, audio reasoning, and cost.

Confirmed

  • Latency: Gemini 3.8 Live averaged a 1.18-second time-to-first-audio on Big Bench Audio, with Extended Thinking (High) at 1.35 seconds—roughly 2.5x faster than the previous generation
  • Agent capability: Gemini 3.8 Live Extended Thinking (High) topped the Tau Voice speech-agent benchmark at 68.6%, nearly double the previous generation's score and ahead of GPT-Live and other rivals
  • Audio reasoning: it scored 97.7% on Big Bench Audio, beating models like Grok Voice Think Fast and second only to Qwen
  • Blind preference: in its debut on the Speech Agent Arena real-time voice blind test, it ranked 2nd with an Elo of 1083, behind only the previous-generation Gemini 3.1 Flash Live
  • Pricing: input audio costs $0.84 per hour, the lowest in its Speech-to-Speech index tier and roughly half the price of the previous-generation Gemini 3.1 Flash Live High (the original post used the phrasing "roughly")

Why it matters

  • The data comes from third-party evaluator Artificial Analysis and spans five dimensions—latency, agent benchmarks, audio reasoning, human blind-preference testing, and cost—painting a fairly complete comparative picture of Gemini 3.8 Live
  • The combination of low latency and low cost signals a further drop in the barrier to real-time voice agent applications, potentially intensifying competition with GPT-Live, Grok, and other voice models