Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS with 100+ language custom voices
GoogleAI · x · 2026-09-23
Google has launched two new text-to-speech models:
- Gemini 3.8 Flash TTS: built for high-fidelity creative production (gaming, audiobooks, podcasts), with from-scratch voice design and line-by-line delivery control, including cues like <laughs> and backchannel interjections like |mhm|.
- Gemini 3.8 Flash-Lite TTS: tuned for cost-efficient scale, auto-adjusting tone and pacing for near real-time voice agents, dubbing, and bulk audio.
Both support custom voices in 100+ languages or 2,000+ production-ready voices, can direct back-and-forth conversations, and generate hours of consistent, glitch-free audio.
Related event: Google launches Gemini 3.8 Flash TTS with instant voice cloning(14 posts)→
More from Multimodal
- A Full Gemini Workflow: Audio Transcription, Kinetic Subtitles, and 60fps Rendering — AI_Andrew · 2026-09-24
- DIY Minimax H3 workflow adds multi-image + audio + video reference inputs — TheNeonGrid · 2026-09-24
- How PixVerse R2 works: dynamic chunks adapt to input, multi-timescale memory keeps worlds consistent — SarahAnnabels · 2026-09-24
- Looking for a model to label each subtitle line with the speaker's identity — dtdisapointingresult · 2026-09-24
- fal enterprise spend quadruples in six months; compute is generative media's binding constraint — isidentical · 2026-09-24
- APOB pairs with Seedance 2.5 to turn AI influencer vlogs into consistent 30-second cinematic stories — aftahi_ai · 2026-09-24