Testing shows AI video audio can't mix two reference tracks simultaneously

Portable_Solar_ZA · reddit · 2026-08-29

A developer testing for the Comfy Sync competition found that current audio generation models cannot take two reference audio tracks that play at the same time: with both a reference voice and reference music, the music either disappears when the narrator speaks or comes out barely audible and garbled.

The working pattern is "1 reference + 1 model-generated": reference voice with model-generated music, or model-generated voice over reference music, both of which layer correctly. He had planned a trailer with custom narrator plus reference music and had to abandon that combination.

Original post →

More from Multimodal

Multimodal channel →