Google Launches Gemini 3.8 Flash TTS with 30-Second Voice Cloning
On September 23, Google DeepMind released two next-generation text-to-speech models: Gemini 3.8 Flash TTS for high-fidelity creative work (games, audiobooks, podcasts) and Gemini 3.8 Flash-Lite TTS for high-volume, low-cost scenarios. Google calls them its most expressive audio generation models with SOTA performance; they are available in AI Studio and Gemini platforms, with generated audio watermarked.
Confirmed
- Core capabilities: design a brand-new voice from a one-sentence description (voice design), or clone an existing voice from 30 seconds of audio; 2000+ production-ready preset voices across 100+ languages.
- Fine-grained control: per-line direction of style, tone, pace and emotion; expressive tags and natural cues such as inserted laughter; voice library filterable by language, accent, gender, pitch and persona; custom voices support persistent voice IDs.
- Positioning: Google frames speech synthesis as a "dynamic creative studio," emphasizing voice as an increasingly central mode of human-computer interaction.
- Performance: per Philschmid, the model tops the Hume voice design leaderboard and ranks first in six languages on Voice Arena; Artificial Analysis places it first in pronunciation robustness but second overall in the Provider Voice Arena blind test, behind Sonic 3.6.
- Pricing: per rseroter's hands-on test, generating 1 minute 18 seconds of audio cost only 2.74 US cents, underscoring Flash-Lite's low-cost positioning.
Unconfirmed
- m1 describes the cloning and director-level control under the name "Gemini 3.8 text-to-speech," slightly at odds with the Flash/Flash-Lite dual-model naming in most posts; official blog naming prevails.
- m11's mention of Gemini 3.8 Live with Extended Thinking (background reasoning during voice conversations) appears to be a separate concurrent update with no stated relation to this TTS launch.
Why it matters
- 30-second voice cloning plus prompt-level voice design dramatically lowers the barrier to custom speech, directly challenging ElevenLabs, Hume AI and other voice vendors—though third-party blind tests suggest the competitive picture is not settled.
- Combining 2000+ production-grade preset voices with Flash-Lite's low-cost tier (roughly 2 cents per minute of audio) lets it serve both creative production and large-scale deployment.
2026-09-22 ~ 2026-09-24 · 27 related posts
- Episode 1: Google Launches Gemini 3.8 Live Models, Topping Speech Quality Index at Half Price(2026-09-15, 29 posts)
- Episode 2: Artificial Analysis: Gemini 3.8 Live tops voice agent benchmarks at record low cost(2026-09-16, 7 posts)
- Episode 3: Google Launches Gemini 3.8 Flash TTS with 30-Second Voice Cloning(2026-09-22, 27 posts)
- Episode 4: DeepMind Unveils Gemini 3.8 TTS with Voice Cloning and Multi-Speaker Dialogue(2026-09-24, 3 posts)
Primary sources
- Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS with 100+ language custom voices — GoogleAI ·
- Google launches Gemini 3.8 Flash and Flash-Lite TTS with 2,000+ voices in 100 languages — OfficialLoganK ·
- Google ships Gemini 3.8 Flash TTS: #1 pronunciation robustness, #2 voice arena at 1263 Elo — ArtificialAnlys ·
- Google ships Gemini 3.8 Live with background reasoning during voice calls — emmanuelvivier · 2026-09-22
- Google DeepMind launches Gemini 3.8 TTS with 30-second voice cloning and consent checks — Google DeepMind · 2026-09-23
- [source] Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS with 100+ language custom voices — GoogleAI · 2026-09-23
- Gemini 3.8 TTS tops Hume Voice Design Benchmark and leads Voice Arena in 6 languages — _philschmid · 2026-09-23
- Google's new TTS and voice design model: clone your voice in 30 seconds, design voices from a prompt — ammaar · 2026-09-23
- [source] Google ships Gemini 3.8 Flash TTS: #1 pronunciation robustness, #2 voice arena at 1263 Elo — ArtificialAnlys · 2026-09-23
- Google Releases Two New TTS Models for Original and Cloned Voices in 100+ Languages — tulseedoshi · 2026-09-24
- Google Ships 3.8 Live and Live Thinking Voice Models Across Gemini and APIs — tulseedoshi · 2026-09-24
- Google Launches Gemini 3.8 Flash TTS, and Why Codec-LM Tokens Make |mhm| Backchannels Possible — prdeepakbabu · 2026-09-24
18 near-duplicate retellings: AI_Andrew · _philschmid · patloeber · ArtificialAnlys · OfficialLoganK · rseroter · OfficialLoganK · kimmonismus · Arindam_1729 · Recoil42 · kastnerkyle · AI_Andrew · giffmana · petrusenko_max · altryne · dioscuri · rseroter · goyalshaliniuk