Gemini 3.8 Flash TTS launches: clone a voice from 30s of audio
kimmonismus · x · 2026-09-24
Philipp Schmid (Google) announces Gemini 3.8 Flash TTS and Flash-Lite TTS. Key capabilities:
- Voice replication: clone a voice from just 30 seconds of audio, or design a new one from a text description
- Voice library: filterable by language, accent, gender, pitch, and persona
- Line-by-line direction: control style and delivery for every line
Gemini reportedly ranks #1 on Hume's Voice Design Benchmark and tops the Voice Arena in 6 languages.
More from Multimodal
- Claude + Thrixel Build an Interactive 3D Room Designer You Can Walk Through — RanaHanocka · 2026-09-24
- Hands-On With Pexo: An AI Video Agent That Builds Full Promos via Chat — HeyAmit_ · 2026-09-24
- Opus 5.5 tested on music: agent skill rewrites melodies into virtuosic showpieces — doodlestein · 2026-09-24
- Atlas Teases Chisel in Beta: Block Out a World and Let AI Bring It to Life — theworldlabs · 2026-09-24
- World Labs releases Atlas, an omni world model native to text, images, video, and 3D — theworldlabs · 2026-09-24
- Higgsfield Genjutsu used to fake an entire 'maxxed' lifestyle on video — floguo · 2026-09-24