Gemini 3.1 Flash TTS Struggles with Voice Consistency in Long-Form Audio
Ok_Coat4453 · reddit · 2026-07-06
When exploring Gemini 3.1 Flash TTS Preview for long-form audiobooks and storytelling, the author found maintaining voice consistency to be the biggest challenge. Even using the same text, voice (Kore), API, and parameters, repeated generations exhibit noticeable differences in timbre, tone, speaking style, and overall vocal characteristics, despite the gender remaining consistent. This becomes particularly evident when stitching multiple clips together. Experiments with different chunk sizes, emotion tags, and natural language acting instructions yielded no significant improvements, prompting the author to seek long-form voiceover solutions from the community.
More from Multimodal
- Midjourney style code share: --sref 2912175708 — tisch_eins · 2026-09-11
- The full prompt-to-3D-game workflow: Hyper3D Rodin MCP plus Codex, no reference image — FellMentKE · 2026-09-11
- Building a 3D landing page with GPT-6 Astra and Hyper3D Rodin MCP, no modeling needed — FellMentKE · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- MiniMax H3 MAX nails cooking anime clips: 15-second curry demo with prompts shared — Hailuo_AI · 2026-09-11
- MiniMax Music Production Toolkit 2.5 for ComfyUI adds full mastering chain — Vivid_Promise1700 · 2026-09-11