Testing shows AI video audio can't mix two reference tracks simultaneously
Portable_Solar_ZA · reddit · 2026-08-29
A developer testing for the Comfy Sync competition found that current audio generation models cannot take two reference audio tracks that play at the same time: with both a reference voice and reference music, the music either disappears when the narrator speaks or comes out barely audible and garbled.
The working pattern is "1 reference + 1 model-generated": reference voice with model-generated music, or model-generated voice over reference music, both of which layer correctly. He had planned a trailer with custom narrator plus reference music and had to abandon that combination.
More from Multimodal
- MiniMax H3 + ComfyUI: Transfer Any Video Pose via ControlNet — Maleficent-Tell-2718 · 2026-08-29
- Prompt Showcase: Ultra-Realistic Indonesian Vacation Video Generation — SimplyAnnisa · 2026-08-29
- Rookie run: 84-shot video pipeline using rented GPU and MiniMax H3 — Legitimate_Bit2775 · 2026-08-29
- Prompt breakdown for hyper-realistic miniature solar village — CurieuxExplorer · 2026-08-29
- Opus 5 generates 3D Pokemon town in 5 hours via sub-agents — anselm · 2026-08-29
- APOB AI Generates Realistic AI Influencer Videos; Seedance 2.5 Tackles 30s Vlogs — aftahi_ai · 2026-08-29