Minimax H3 voice cloning proves inconsistent; user shares full ComfyUI workflow
Elzair · reddit · 2026-10-11
A user tried H3 Minimax's audio-reference feature to keep a character's voice consistent across clips, with mixed results. They shared the full pipeline: extracting a reference clip, feeding it back for generation, and comparing against an externally generated VoxCPM2 voice, along with the complete ComfyUI node prompt (subjectdefinitions, retentionanalysis, etc.). They ask the community whether it's a bad seed or wrong usage, and whether voices should be generated externally and plugged in.
More from Multimodal
- Seedance 2.5 Video Demo: Fixing an Old Lamp with a Single Prompt — SimplyAnnisa · 2026-10-11
- Step 5 Preview generates a 30-second motion clip in one shot via Hermes agent — Teknium · 2026-10-11
- Tencent's open-source Hunyuan3D-2 turns one image into textured 3D models, 15k stars — Promptmethus · 2026-10-11
- Open-source music model YuE2 runs locally on a MacBook Pro, rivaling Suno quality — vista8 · 2026-10-11
- Two prompts + Claude made a Monet-style MV for Jay Chou's 'Qi Li Xiang' — AlchainHust · 2026-10-11
- Underdog launches on-device image gen powered by Qwen models, photos never leave your computer — Scobleizer · 2026-10-11