Gemini 3.1 Flash TTS Struggles with Voice Consistency in Long-Form Audio

Ok_Coat4453 · reddit · 2026-07-06

When exploring Gemini 3.1 Flash TTS Preview for long-form audiobooks and storytelling, the author found maintaining voice consistency to be the biggest challenge. Even using the same text, voice (Kore), API, and parameters, repeated generations exhibit noticeable differences in timbre, tone, speaking style, and overall vocal characteristics, despite the gender remaining consistent. This becomes particularly evident when stitching multiple clips together. Experiments with different chunk sizes, emotion tags, and natural language acting instructions yielded no significant improvements, prompting the author to seek long-form voiceover solutions from the community.

Original post →

More from Multimodal

Multimodal channel →