Google Launches Gemini 3.8 Flash TTS, and Why Codec-LM Tokens Make |mhm| Backchannels Possible

prdeepakbabu · x · 2026-09-24

Google launched Gemini 3.8 Flash TTS and Flash-Lite TTS, its most expressive audio models: custom voices in 100+ languages or 2,000+ presets, line-by-line delivery direction, and natural conversational cues like <laughs> and active-listening interjections like |mhm|, with hours-long consistent generation.

Developer Deepak Babu Piskala highlights the real technical story: a backchannel must be emitted without claiming the turn, and mel-spectrogram TTS has no token to represent that — discrete codec-LM TTS makes it representable. He points to Chapter 7 of his book Building Speech AI, a 320+ page practitioner's guide with runnable code, notebooks, and CLI scripts spanning speech representation, recognition (Whisper, wav2vec 2.0, Conformer), and production voice-system trade-offs.

Related event: Google Launches Gemini 3.8 Flash TTS with 30-Second Voice Cloning(27 posts)→

Original post →

More from Multimodal

Multimodal channel →