Google DeepMind launches Gemini 3.8 TTS with 30-second voice cloning and consent checks
Google DeepMind · youtube · 2026-09-23
Google DeepMind officially launched its Gemini 3.8 text-to-speech models:
- Custom voices: design entirely new vocal personas via natural-language prompts for gaming, audiobooks, or dual-speaker podcasts, directing each line with acting cues, pacing, back-channeling, and dialect shifts.
- Voice cloning: recreate a consistent adult vocal profile from just a 30-second sample you have rights to use.
- Safety rails: built-in consent verification, SynthID watermarking, and C2PA credentials.
The models are available now via the Gemini Audio lineup.
Related event: Google Launches Gemini 3.8 Flash and Flash-Lite TTS Models(13 posts)→
More from Multimodal
- SupraLabs' from-scratch text-to-image model Supra2-IMG trends on Hugging Face — SupraLabs · 2026-09-24
- Hum-to-song app yue2-hum-to-song trends on Hugging Face — Mothersuperior · 2026-09-24
- Grok 4.7 turns a single image into a full 3D jet model inside Grok Build — XFreeze · 2026-09-24
- MiniMax H3 on Spectrum Runs Each Step Twice, Killing Performance — Glittering-Cold-2981 · 2026-09-24
- Google's TTS Update Lets Prompts Customize Voice Tone, First Step Toward Personalized Agents — kimmonismus · 2026-09-24
- Comfy Router launches: one API to route image, video, 3D and audio models across providers — PurzBeats · 2026-09-24