MiniMax H3 multi-subject video bug: Voices bleed between different reference characters
Wide-Researcher583 · reddit · 2026-08-10
Users have discovered a significant flaw when using the Ref2VA model of MiniMax H3: voices 'bleed' between different subjects when processing separate reference images containing multiple characters.
Specifically, while the visual identities of different subjects are perfectly preserved, the audio branch blends the voices or one subject's voice identity leaks into another. Users have tried various methods such as changing seeds, shortening clips, and explicitly defining subjects, but no reliable solution has been found yet.
More from Models
- Testing MiniMax Video Model: Generating with 50+ Reference Images — No_Damage_8420 · 2026-08-10
- Dev Test: DeepSeek Often Outperforms Sol — yacineMTB · 2026-08-10
- NVIDIA Shares DGX Spark Guide for Local LLM Deployment — lifebypixels · 2026-08-10
- 35B Model Beats Giants? Community Suspects Benchmark Contamination — LegacyRemaster · 2026-08-10
- Pichai Admits Google's Best Model Is Shelved: The Bottleneck is Serving Cost, Not Intelligence — r0ck3t23 · 2026-08-10
- Pushing OpenAI Codex to the Limit: 4-Hour PR Grind Burns 9% of Weekly Credits — andimarafioti · 2026-08-10