Developer Seeks Fix for Gibberish in Long-Take Voice Cloning with WAV Samples
davekilljoy · reddit · 2026-08-09
A developer is struggling to achieve consistent voice cloning using a 10-second WAV reference. When applying this to 10-15 second long single-camera shots, the character outputs gibberish 90% of the time.
Currently, the only workaround is using multiple cuts in the video. The author is asking the community for a reliable way to generate cloned audio for long-running clips without relying on frequent cuts.
More from Multimodal
- Grok Image 2.0 Tested: Fancy Segmentation Enables Precise Image Editing — mark_k · 2026-08-09
- RTX 5060Ti Test: Generating Cinematic Action Video with Minimax H3 Turbo — AppointmentOk972 · 2026-08-09
- Running Minimax H3 Cinematic Action Video Workflow on a Single 5060Ti — AppointmentOk972 · 2026-08-09
- Seedance 2.5 Stuns Users with Realism Rivaling Bollywood BTS — taherdhanera · 2026-08-09
- MiniMax H3 Video Acceleration: 4-Step LoRA & Decoupled Audio/Video Scheduling — wjc_5 · 2026-08-09
- Running MiniMax H3 on 12GB VRAM: RTX 5090 Generates 1080p Video in 13 Mins — Fun_Firefighter_7785 · 2026-08-09