MiniMax i2v Workflow: Combining fl2va and ref2va for Voice Cloning
spiderofmars · reddit · 2026-08-14
The author shares an accidental discovery of achieving voice cloning in MiniMax image-to-video workflows using the fl2va and ref2va models.
- Workflow Setup: Input a first-frame reference image and an audio sample into the ref2va node to clone the voice for new dialogue.
- Model Comparison: The fl2va model offers better sound quality and natural motion, while ref2va has a thinner sound but adds unprompted camera movements (e.g., simulating a driving perspective).
- Evaluation: The author notes a 95% voice similarity, though cloned voices seem to lack the emotional inflection of randomly generated ones.
More from Multimodal
- Meta/Oxford study: multimodal models need only 5% image-generation data, language training is key — rohanpaul_ai · 2026-08-14
- Is MiniMax H3 Ref2Vid Model Only Available via Cloud API? — SensitiveUse7864 · 2026-08-14
- 18-sec One-Shot! Seedance 2.5 Maldives Resort Video Prompt — techhalla · 2026-08-14
- DotSight AI Releases Open-Source 280B Multimodal Model with Apache 2.0 License — Xianbao_QIAN · 2026-08-14
- Wan 3.0 official prompts revealed: 64 prompts as 30-second shot scripts — Tricky_Algae2625 · 2026-08-14
- AI horrorcore diss track 'Syko Sam' released, showcasing AI music creation — JonMillaTheKilla · 2026-08-14