Google launches Gemini 3.8 TTS: 2,000+ voices, voice cloning, 1m18s audio for 2.74 cents
rseroter · x · 2026-09-24
Google shipped two new text-to-speech models, gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts, focused on low cost and multi-speaker conversation generation.
- Ships with 2,000+ preset voices plus custom voice cloning from a 30-second audio sample
- The API makes it easy to script full multi-character conversations, each with distinct voices and style instructions
- Simon Willison vibe-coded a bring-your-own-key TTS Playground with GPT-6 Astra, leveraging the Gemini API's open CORS policy; it supports single-voice narration and multi-speaker dialogues with shareable URLs
- Benchmark from his demo: 20 seconds to generate 1m 18s of audio on Flash TTS at a cost of 2.74 cents; the two-pelicans-debating script was written by Claude 4.5 Opus
Related event: Google Launches Gemini 3.8 Flash TTS with 30-Second Voice Cloning(27 posts)→
More from Multimodal
- FineVision, the 17M-image open VLM dataset from 200+ sources, accepted to NeurIPS — andimarafioti · 2026-09-25
- MiniMax H3 seems overtrained on smiles: 'bored caterpillar' video prompt keeps breaking immersion — episodex86 · 2026-09-25
- Pose Blueprint: A Browser-Based 3D Pose Editor for ComfyUI and ControlNet — OkConfusion6667 · 2026-09-25
- Reddit User Explores AI Art With Only Steps, CFG and Denoise Tweaks, No LoRAs — Extreme_Nice · 2026-09-25
- A sub-$20 LoRA makes Qwen-Image 2.1 rotate transparent objects with a prompt — ben_burtenshaw · 2026-09-25
- Lingbot World v2 runs at 60 FPS, hinting world models could reshape game dev — bingxu_ · 2026-09-25