Google launches Gemini 3.8 Flash TTS, tops Hume's Voice Design Benchmark
_philschmid · x · 2026-09-23
Google released Gemini 3.8 Flash TTS and Flash-Lite TTS: clone a voice from 30s of audio or create one from a text description, with 2,000+ ready voices across 100+ languages, line-by-line direction of style and natural cues like <laughs>, and hours of consistent generation. The models rank #1 on Hume's Voice Design Benchmark and top Voice Arena in 6 languages. Blog, prompt guide, and migration guide from 3.1 are available.
More from Multimodal
- A Full Gemini Workflow: Audio Transcription, Kinetic Subtitles, and 60fps Rendering — AI_Andrew · 2026-09-24
- DIY Minimax H3 workflow adds multi-image + audio + video reference inputs — TheNeonGrid · 2026-09-24
- How PixVerse R2 works: dynamic chunks adapt to input, multi-timescale memory keeps worlds consistent — SarahAnnabels · 2026-09-24
- Looking for a model to label each subtitle line with the speaker's identity — dtdisapointingresult · 2026-09-24
- fal enterprise spend quadruples in six months; compute is generative media's binding constraint — isidentical · 2026-09-24
- APOB pairs with Seedance 2.5 to turn AI influencer vlogs into consistent 30-second cinematic stories — aftahi_ai · 2026-09-24