Google launches Gemini 3.8 Flash TTS with 100+ languages and 30-second voice cloning
rseroter · x · 2026-09-23
Google officially announced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, positioning speech generation as a dynamic creative studio:
- Design bespoke voices from scratch via prompts across 100+ languages
- Control tone, accents, and pacing line by line with script cues
- Stage two-speaker dialogues with natural turn-taking from a single script
- Script nonverbal cues like laughs, sighs, and listening interjections
- Recreate consistent vocal profiles from a 30-second sample (with usage rights)
- Access 2,000+ production-ready voices
All audio is watermarked with SynthID for detectability.
More from Multimodal
- ComfyUI Launches Comfy Router: One API for Image, Video, 3D and Audio Models — cpaik · 2026-09-24
- Turning any image into a thousand stories with Midjourney + Seedance 2.5 on Runway — umesh_ai · 2026-09-24
- Google announces Gemini 3.8 Flash TTS and Flash-Lite TTS models — Recoil42 · 2026-09-24
- A Full Gemini Workflow: Audio Transcription, Kinetic Subtitles, and 60fps Rendering — AI_Andrew · 2026-09-24
- DIY Minimax H3 workflow adds multi-image + audio + video reference inputs — TheNeonGrid · 2026-09-24
- How PixVerse R2 works: dynamic chunks adapt to input, multi-timescale memory keeps worlds consistent — SarahAnnabels · 2026-09-24