Ref2VA Tested: Cloning Audio and Video with 3 Reference Images
rjay7979 · reddit · 2026-08-05
A developer showcased a practical test of the Ref2VA model for audio-driven video generation. In the demo, the user provided the model with 3 reference images of different subjects and an audio clip of Gilbert Gottfried. The model successfully generated a video based on these inputs, with the author noting they were testing if the model natively recognizes the celebrity's voice.
More from Multimodal
- Precise 3D Spatial Localization of Photos Unlocks New AI Possibilities — bilawalsidhu · 2026-08-05
- Live Wallpaper Comparison: Wan 2.2 vs LTX 2.3 vs Minimax H3 — aziib · 2026-08-05
- Developer Uses Hailuo AI's New Model to Generate Original Music — jordanbayne · 2026-08-05
- Live Wallpaper Comparison: Wan 2.2 vs LTX 2.3 vs Minimax H3 — aziib · 2026-08-05
- Experimental AI Film Trailer Showcase — 5oclockgoldenhour · 2026-08-05
- Training Krea 2 LoRAs: Can We Reuse Old Tag-Based Datasets? — poliranter · 2026-08-05