Google Launches Gemini 3.8 Flash and Flash-Lite TTS with 30-Second Voice Cloning Across 100+ Languages
On September 23, Google released two next-generation text-to-speech models: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. Flash TTS targets high-fidelity creative production (games, audiobooks, podcasts), while Flash-Lite TTS is built for large-scale, low-cost use cases. Both cover 100+ languages and top speech design benchmarks, marking a major TTS upgrade for Google—one worth watching for its impact on audio content and game voiceover production pipelines.
Confirmed
- Supports voice cloning from 30 seconds of audio, or designing an entirely new voice persona from scratch via a single natural-language description, suited for games, audiobooks, and two-host podcasts
- Offers a voice library filterable by language, accent, gender, pitch, and persona, with 2000+ preset voices
- Enables line-by-line director-style control over performance cues, style, and pacing, including natural markers like <laughs> as well as backchannels and dialect switching
- According to Philschmid, the model tops the Hume speech design leaderboard and ranks first across all six languages evaluated in Voice Arena
Why it matters
- Dropping the voice cloning barrier to 30 seconds of audio, combined with line-level expressiveness control, dramatically cuts production costs for game voiceovers, audiobooks, and podcasts
- 100+ languages and 2000+ preset voices cover a wide range of localization and bulk synthesis needs, with the Flash-Lite variant further lowering costs for large-scale deployment
- Its lead on third-party speech design benchmarks signals a strong technical position against competitors like Hume, potentially sparking a new round of competition in speech generation
2026-09-23 ~ 2026-09-24 · 14 related posts
Primary sources
- Google DeepMind launches Gemini 3.8 TTS with 30-second voice cloning and consent checks — Google DeepMind ·
- Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS with 100+ language custom voices — GoogleAI ·
- Gemini 3.8 TTS tops Hume Voice Design Benchmark and leads Voice Arena in 6 languages — _philschmid ·
- [source] Google DeepMind launches Gemini 3.8 TTS with 30-second voice cloning and consent checks — Google DeepMind · 2026-09-23
- [source] Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS with 100+ language custom voices — GoogleAI · 2026-09-23
- [source] Gemini 3.8 TTS tops Hume Voice Design Benchmark and leads Voice Arena in 6 languages — _philschmid · 2026-09-23
- Google's new TTS and voice design model: clone your voice in 30 seconds, design voices from a prompt — ammaar · 2026-09-23
- Google ships Gemini 3.8 Flash TTS: #1 pronunciation robustness, #2 voice arena at 1263 Elo — ArtificialAnlys · 2026-09-23
9 near-duplicate retellings: AI_Andrew · _philschmid · patloeber · ArtificialAnlys · OfficialLoganK · rseroter · OfficialLoganK · kimmonismus · Arindam_1729