Google Launches Gemini 3.8 Flash TTS, and Why Codec-LM Tokens Make |mhm| Backchannels Possible
prdeepakbabu · x · 2026-09-24
Google launched Gemini 3.8 Flash TTS and Flash-Lite TTS, its most expressive audio models: custom voices in 100+ languages or 2,000+ presets, line-by-line delivery direction, and natural conversational cues like <laughs> and active-listening interjections like |mhm|, with hours-long consistent generation.
Developer Deepak Babu Piskala highlights the real technical story: a backchannel must be emitted without claiming the turn, and mel-spectrogram TTS has no token to represent that — discrete codec-LM TTS makes it representable. He points to Chapter 7 of his book Building Speech AI, a 320+ page practitioner's guide with runnable code, notebooks, and CLI scripts spanning speech representation, recognition (Whisper, wav2vec 2.0, Conformer), and production voice-system trade-offs.
Related event: Google Launches Gemini 3.8 Flash TTS with 30-Second Voice Cloning(27 posts)→
More from Multimodal
- FineVision, the 17M-image open VLM dataset from 200+ sources, accepted to NeurIPS — andimarafioti · 2026-09-25
- MiniMax H3 seems overtrained on smiles: 'bored caterpillar' video prompt keeps breaking immersion — episodex86 · 2026-09-25
- Pose Blueprint: A Browser-Based 3D Pose Editor for ComfyUI and ControlNet — OkConfusion6667 · 2026-09-25
- Reddit User Explores AI Art With Only Steps, CFG and Denoise Tweaks, No LoRAs — Extreme_Nice · 2026-09-25
- A sub-$20 LoRA makes Qwen-Image 2.1 rotate transparent objects with a prompt — ben_burtenshaw · 2026-09-25
- Lingbot World v2 runs at 60 FPS, hinting world models could reshape game dev — bingxu_ · 2026-09-25