Ref2VA Tested: Cloning Audio and Video with 3 Reference Images

rjay7979 · reddit · 2026-08-05

A developer showcased a practical test of the Ref2VA model for audio-driven video generation. In the demo, the user provided the model with 3 reference images of different subjects and an audio clip of Gilbert Gottfried. The model successfully generated a video based on these inputs, with the author noting they were testing if the model natively recognizes the celebrity's voice.

Original post →

More from Multimodal

Multimodal channel →