Gemini 3.8 TTS tops Hume Voice Design Benchmark and leads Voice Arena in 6 languages
_philschmid · x · 2026-09-23
Philschmid shares key details on Gemini 3.8 Flash TTS and Flash-Lite TTS:
- Clone a voice from 30 seconds of audio or describe one in a sentence; voice library filterable by language, accent, gender, pitch, persona.
- Line-by-line direction: style, <laugh>, <sigh>, backchannels, and 2-speaker scenes.
- Benchmarks: #1 on Hume Voice Design, #1 and #2 on Hume Quality, top of Voice Arena in 6 languages.
- Safety: consent check on voice replication and SynthID watermarking on all clips.
Voices are created once and reusable in any TTS call, available via the Gemini API and Google AI Studio.
Related event: Google launches Gemini 3.8 Flash TTS with instant voice cloning(14 posts)→
More from Multimodal
- A Full Gemini Workflow: Audio Transcription, Kinetic Subtitles, and 60fps Rendering — AI_Andrew · 2026-09-24
- DIY Minimax H3 workflow adds multi-image + audio + video reference inputs — TheNeonGrid · 2026-09-24
- How PixVerse R2 works: dynamic chunks adapt to input, multi-timescale memory keeps worlds consistent — SarahAnnabels · 2026-09-24
- Looking for a model to label each subtitle line with the speaker's identity — dtdisapointingresult · 2026-09-24
- fal enterprise spend quadruples in six months; compute is generative media's binding constraint — isidentical · 2026-09-24
- APOB pairs with Seedance 2.5 to turn AI influencer vlogs into consistent 30-second cinematic stories — aftahi_ai · 2026-09-24